To content
Lecture Series

AI Colloquium

The AI Colloquium is a series of lectures dedicated to cutting-edge research in the field of machine learning and artificial intelligence, coorganized by the Lamarr Institute for Machine Learning and Artificial Intelligence (Lamarr Institute), the Research Center Trustworthy Data Science and Security (RC Trust), and the Center for Data Science & Simulation at TU Dortmund University (DoDas).

Programme

Distinguished researchers deliver captivating lectures followed by vibrant discussions. However, unlike traditional colloquia, the AI Colloquium prioritizes interactive dialogue, fostering international collaboration. Conducted primarily in English, these 90-minute sessions feature hour-long lectures and 30-minute Q&A sessions. Join every Thursday at 10 AM c.t. for a stimulating exploration of cutting-edge topics. Whether in-person at our Lecture Room on Fraunhofer Strasse 25 or via Zoom, our hybrid format ensures accessibility for all.

Day (usually) Thursday
Start and end time 10 AM c.t. - 12 AM
Duration of Presentation 60 Minutes
Location (usually) Lecture Room 303
3. Floor
Fraunhofer Strasse 25
Dortmund

Upcomming Events

ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules

Start: End: Location: Joseph von Fraunhofer Strassse 25 | 3-303 - Conference Room (Lamarr/RC Trust Dortmund)
Event type:
  • Resource-aware ML
Jonas Landsgesell (University of Stuttgart), Pascall Knoll (University of Stuttgart), Tizian Wenzel (Ludwig Maximilian University of Munich, Munich Center for Machine Learning)

Abstract: While select tabular foundation models (TFMs), such as TabPFN and TabICL, are capable of producing full predictive distributions, prevailing benchmarks often rely exclusively on point-estimate metrics like RMSE. This focus rewards conditional mean accuracy while potentially ignoring the valuable distributional information these models are intended to capture.
To address this evaluation gap, we introduce ScoringBench, an open and extensible benchmark designed to evaluate tabular regression models using a comprehensive suite of proper scoring rules alongside standard point metrics.
Key Findings and Implications:

  • Metric Sensitivity: Across diverse model families, we observe that rankings shift substantially depending on the chosen metric. Models that perform well on point metrics often rank poorly when evaluated through probabilistic lenses.
  • Inductive Biases: Different proper scoring rules impose distinct inductive biases during training. Even when theoretically minimized by the true distribution, the choice of loss function directly influences the model’s behavior.
  • Fine-tuning Potential: We demonstrate that fine-tuning with scoring rules that were unseen during the pretraining phase can lead to improvements in the corresponding metrics.


These results suggest that the choice of metric is a critical modeling decision. We argue that benchmarks should prioritize the reporting of distributional metrics, and that consideration should be given to how a model's training objective aligns with the specific requirements of the downstream decision problem. Ultimately, choosing the loss function is an integral part of choosing the model itself.

Archiv

Past Events

ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules

Start: End: Location: Joseph von Fraunhofer Strassse 25 | 3-303 - Conference Room (Lamarr/RC Trust Dortmund)
Event type:
  • Resource-aware ML
Jonas Landsgesell (University of Stuttgart), Pascall Knoll (University of Stuttgart), Tizian Wenzel (Ludwig Maximilian University of Munich, Munich Center for Machine Learning)

Abstract: While select tabular foundation models (TFMs), such as TabPFN and TabICL, are capable of producing full predictive distributions, prevailing benchmarks often rely exclusively on point-estimate metrics like RMSE. This focus rewards conditional mean accuracy while potentially ignoring the valuable distributional information these models are intended to capture.
To address this evaluation gap, we introduce ScoringBench, an open and extensible benchmark designed to evaluate tabular regression models using a comprehensive suite of proper scoring rules alongside standard point metrics.
Key Findings and Implications:

  • Metric Sensitivity: Across diverse model families, we observe that rankings shift substantially depending on the chosen metric. Models that perform well on point metrics often rank poorly when evaluated through probabilistic lenses.
  • Inductive Biases: Different proper scoring rules impose distinct inductive biases during training. Even when theoretically minimized by the true distribution, the choice of loss function directly influences the model’s behavior.
  • Fine-tuning Potential: We demonstrate that fine-tuning with scoring rules that were unseen during the pretraining phase can lead to improvements in the corresponding metrics.


These results suggest that the choice of metric is a critical modeling decision. We argue that benchmarks should prioritize the reporting of distributional metrics, and that consideration should be given to how a model's training objective aligns with the specific requirements of the downstream decision problem. Ultimately, choosing the loss function is an integral part of choosing the model itself.

ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules

Start: End: Location: Joseph von Fraunhofer Strassse 25 | 3-303 - Conference Room (Lamarr/RC Trust Dortmund)
Event type:
  • Resource-aware ML
Jonas Landsgesell (University of Stuttgart), Pascall Knoll (University of Stuttgart), Tizian Wenzel (Ludwig Maximilian University of Munich, Munich Center for Machine Learning)

Abstract: While select tabular foundation models (TFMs), such as TabPFN and TabICL, are capable of producing full predictive distributions, prevailing benchmarks often rely exclusively on point-estimate metrics like RMSE. This focus rewards conditional mean accuracy while potentially ignoring the valuable distributional information these models are intended to capture.
To address this evaluation gap, we introduce ScoringBench, an open and extensible benchmark designed to evaluate tabular regression models using a comprehensive suite of proper scoring rules alongside standard point metrics.
Key Findings and Implications:

  • Metric Sensitivity: Across diverse model families, we observe that rankings shift substantially depending on the chosen metric. Models that perform well on point metrics often rank poorly when evaluated through probabilistic lenses.
  • Inductive Biases: Different proper scoring rules impose distinct inductive biases during training. Even when theoretically minimized by the true distribution, the choice of loss function directly influences the model’s behavior.
  • Fine-tuning Potential: We demonstrate that fine-tuning with scoring rules that were unseen during the pretraining phase can lead to improvements in the corresponding metrics.


These results suggest that the choice of metric is a critical modeling decision. We argue that benchmarks should prioritize the reporting of distributional metrics, and that consideration should be given to how a model's training objective aligns with the specific requirements of the downstream decision problem. Ultimately, choosing the loss function is an integral part of choosing the model itself.