AI Colloquium
The AI Colloquium is a series of lectures dedicated to cutting-edge research in the field of machine learning and artificial intelligence, coorganized by the Lamarr Institute for Machine Learning and Artificial Intelligence (Lamarr Institute), the Research Center Trustworthy Data Science and Security (RC Trust), and the Center for Data Science & Simulation at TU Dortmund University (DoDas).
Programme
Distinguished researchers deliver captivating lectures followed by vibrant discussions. However, unlike traditional colloquia, the AI Colloquium prioritizes interactive dialogue, fostering international collaboration. Conducted primarily in English, these 90-minute sessions feature hour-long lectures and 30-minute Q&A sessions. Join every Thursday at 10 AM c.t. for a stimulating exploration of cutting-edge topics. Whether in-person at our Lecture Room on Fraunhofer Strasse 25 or via Zoom, our hybrid format ensures accessibility for all.
| Day (usually) | Thursday |
| Start and end time | 10 AM c.t. - 12 AM |
| Duration of Presentation | 60 Minutes |
| Location (usually) | Lecture Room 303 3. Floor Fraunhofer Strasse 25 Dortmund |
Upcomming Events
ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules
- Resource-aware ML
Abstract: While select tabular foundation models (TFMs), such as TabPFN and TabICL, are capable of producing full predictive distributions, prevailing benchmarks often rely exclusively on point-estimate metrics like RMSE. This focus rewards conditional mean accuracy while potentially ignoring the valuable distributional information these models are intended to capture.
To address this evaluation gap, we introduce ScoringBench, an open and extensible benchmark designed to evaluate tabular regression models using a comprehensive suite of proper scoring rules alongside standard point metrics.
Key Findings and Implications:
- Metric Sensitivity: Across diverse model families, we observe that rankings shift substantially depending on the chosen metric. Models that perform well on point metrics often rank poorly when evaluated through probabilistic lenses.
- Inductive Biases: Different proper scoring rules impose distinct inductive biases during training. Even when theoretically minimized by the true distribution, the choice of loss function directly influences the model’s behavior.
- Fine-tuning Potential: We demonstrate that fine-tuning with scoring rules that were unseen during the pretraining phase can lead to improvements in the corresponding metrics.
These results suggest that the choice of metric is a critical modeling decision. We argue that benchmarks should prioritize the reporting of distributional metrics, and that consideration should be given to how a model's training objective aligns with the specific requirements of the downstream decision problem. Ultimately, choosing the loss function is an integral part of choosing the model itself.
Past Events
ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules
- Resource-aware ML
Abstract: While select tabular foundation models (TFMs), such as TabPFN and TabICL, are capable of producing full predictive distributions, prevailing benchmarks often rely exclusively on point-estimate metrics like RMSE. This focus rewards conditional mean accuracy while potentially ignoring the valuable distributional information these models are intended to capture.
To address this evaluation gap, we introduce ScoringBench, an open and extensible benchmark designed to evaluate tabular regression models using a comprehensive suite of proper scoring rules alongside standard point metrics.
Key Findings and Implications:
- Metric Sensitivity: Across diverse model families, we observe that rankings shift substantially depending on the chosen metric. Models that perform well on point metrics often rank poorly when evaluated through probabilistic lenses.
- Inductive Biases: Different proper scoring rules impose distinct inductive biases during training. Even when theoretically minimized by the true distribution, the choice of loss function directly influences the model’s behavior.
- Fine-tuning Potential: We demonstrate that fine-tuning with scoring rules that were unseen during the pretraining phase can lead to improvements in the corresponding metrics.
These results suggest that the choice of metric is a critical modeling decision. We argue that benchmarks should prioritize the reporting of distributional metrics, and that consideration should be given to how a model's training objective aligns with the specific requirements of the downstream decision problem. Ultimately, choosing the loss function is an integral part of choosing the model itself.
ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules
- Resource-aware ML
Abstract: While select tabular foundation models (TFMs), such as TabPFN and TabICL, are capable of producing full predictive distributions, prevailing benchmarks often rely exclusively on point-estimate metrics like RMSE. This focus rewards conditional mean accuracy while potentially ignoring the valuable distributional information these models are intended to capture.
To address this evaluation gap, we introduce ScoringBench, an open and extensible benchmark designed to evaluate tabular regression models using a comprehensive suite of proper scoring rules alongside standard point metrics.
Key Findings and Implications:
- Metric Sensitivity: Across diverse model families, we observe that rankings shift substantially depending on the chosen metric. Models that perform well on point metrics often rank poorly when evaluated through probabilistic lenses.
- Inductive Biases: Different proper scoring rules impose distinct inductive biases during training. Even when theoretically minimized by the true distribution, the choice of loss function directly influences the model’s behavior.
- Fine-tuning Potential: We demonstrate that fine-tuning with scoring rules that were unseen during the pretraining phase can lead to improvements in the corresponding metrics.
These results suggest that the choice of metric is a critical modeling decision. We argue that benchmarks should prioritize the reporting of distributional metrics, and that consideration should be given to how a model's training objective aligns with the specific requirements of the downstream decision problem. Ultimately, choosing the loss function is an integral part of choosing the model itself.




