Mátyás Vincze
[ˈmaːcaːʃ]PhD researcher / language agents / reinforcement learning
I am a PhD researcher at the University of Trento , conducting research at Fondazione Bruno Kessler with Bruno Lepri and Giovanni Iacca . I am currently visiting Michiel Bakker 's group at MIT .
I work on post-training for long-horizon language agents, combining distillation, multi-agent feedback, and reinforcement learning with verifiable rewards. My earlier work on interpretable reinforcement learning led to SMoSE , a sparse mixture of shallow policy experts.
Updates
- Visiting Michiel Bakker 's group at MIT to work on post-training for long-horizon language agents 🚀
- ISCRA compute grant from CINECA — more A100s go brrr.
- Driving license secured ✅
- ISCRA compute grant from CINECA — A100s go brrr, again.
- Attended the ALPS 2025 winter school in Aussois, France.
- Presented SMoSE at AAAI 2025 in Philadelphia.
- SMoSE accepted to AAAI 2025 .
- Placed 1st in the Interpretable Control Competition at GECCO 2024.
- ISCRA compute grant from CINECA — A100s go brrr.
Selected publications
see all- A benchmark of expert-level academic questions to assess AI capabilities (Humanity’s Last Exam)Center for AI Safety, Scale AI, HLE Contributors ConsortiumContributor to the expert-level AI benchmark published in Nature.
- SMoSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control TasksM. Vincze, L. Ferrarotti, L. L. Custode, B. Lepri, G. IaccaFirst-author work on sparse, interpretable reinforcement-learning policies with only 108–672 active actor parameters.@ AAAI · 2025