About

I’m a PhD candidate in Computer Science at the National University of Singapore. My research focuses on statistical machine learning, prediction under uncertainty, and robust decision-making. My work combines theory with experiments and has appeared at NeurIPS, ICML, ICLR, and AISTATS.

I expect to complete my PhD in January 2027 and am looking for full-time roles and internships in quantitative research or machine learning research. If you’re hiring or interested in working together, I’d be glad to hear from you—please email me.

Research

I study how to make the most of several models or experts—a question that connects learning-to-defer, model routing, and orchestration. Given an input $x$, the goal is to learn which model $f_i$ is best suited to handle it, subject to the task’s constraints. My work combines theory and algorithms to understand and improve this choice.

Routing an input to one of several models A time series and its context form input x. A learned router selects the second of three available experts, a tree model, to produce a prediction. Selection considers predictive performance and constraints such as cost or latency. The highlighted route is illustrative. Input x History + context Router r(x) Goal + constraints Models / experts f₁ Linear model f₂ Tree model f₃ Neural model Prediction ŷ Routing an input to one of several models Input x passes through a router, which selects a tree model from three available experts. Selection considers predictive performance and constraints such as cost or latency. The selected model produces the prediction. Input x History + context Goal + constraints Router r(x) f₁ Linear f₂ Tree model f₃ Neural ŷ Prediction: f₂(x)
One input, one selected expert: ŷ = fr(x)(x).

Applications range from language models and time-series forecasting to computer vision. More generally, the approach can be used wherever several models or experts are available to query.

Keywords: statistical machine learning, model routing, model orchestration, learning-to-defer, online learning, time series, large language models.

Publications

2026

  1. Beyond Augmented-Action Surrogates for Multi-Expert Learning-to-Defer. Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. NeurIPS26. arXiv:2604.09414
  2. A Query Is Not a Commitment: Learning to Correct Expert Answers in Online Deferral. Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. arXiv submission pending.
  3. Consistent Learning-to-Defer with Expert-Conditioned Advice. Yannis Montreuil, Leina Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. arXiv:2603.14324
  4. Learning to Defer in Non-Stationary Time Series via Switching State-Space Models. Yannis Montreuil*, Letian Yu*, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. arXiv:2601.22538
  5. Why Ask One When You Can Ask k? Learning-to-Defer to the Top-k Experts. Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. ICLR 2026. arXiv:2504.12988.
  6. Online Learning-to-Defer with Varying Experts. Yannis Montreuil*, Duy Dang Hoang*, Maxime Meyer*, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. AISTATS 2026. arXiv:2605.12340.
  7. Adversarial Robustness in One-Stage Learning-to-Defer. Yannis Montreuil*, Letian Yu*, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. AISTATS 2026. arXiv:2510.10988.
  8. Optimal Query Allocation in Extractive QA with LLMs: A Learning-to-Defer Framework with Theoretical Guarantees. Yannis Montreuil*, Yeo Shu Heng*, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. AISTATS 2026. arXiv:2410.15761.
  9. Towards Robust Human–AI Decision-Making via Learning-to-Defer. Yannis Montreuil. AAAI-26 Doctoral Consortium.

2025

  1. Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and Guarantees. Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. ICML 2025. arXiv:2502.01027.
  2. A Two-Stage Learning-to-Defer Approach for Multi-Task Learning. Yannis Montreuil*, Yeo Shu Heng*, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi. ICML 2025. arXiv:2410.15729.

* indicates equal contribution. Abstracts and details are on the learning-to-defer publications page.