2026
- EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot PlanningYichao Liang, Amber Li, Dat Nguyen*, Emily Bunnapradist*, Michelangelo Naim*, Sreela Kodali*, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, and Kevin EllisSep 2026
A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We present EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms. The learned programs can introduce new forces, constraints, and hidden state, and Bayesian inference estimates their parameters and states from noisy observations. The resulting model lets the agent predict the outcomes of actions, choose informative experiments, and revise its hypotheses when predictions fail. Across five simulated domains, EMPIRIC learns interpretable, reusable models, and solves more tasks with fewer environment interactions than all three baselines. On a physical robot, it learns wind forces and domino masses to solve a manipulation task. Website and code: https://yichao-liang.github.io/empiric
@misc{liang2026empiric, title = {EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning}, author = {Liang, Yichao and Li, Amber and Nguyen, Dat and Bunnapradist, Emily and Naim, Michelangelo and Kodali, Sreela and Merler, Matteo and Li, Bowen and Gopinathan, Kiran and Liu, Yiyun and Pimpalkhare, Nikhil and Tenenbaum, Joshua B. and Weller, Adrian and Tavares, Zenna and Silver, Tom and Ellis, Kevin}, year = {2026}, month = sep, cv_date = {2026-09-28}, eprint = {2609.35047}, archiveprefix = {arXiv}, primaryclass = {cs.RO}, url = {https://arxiv.org/abs/2609.35047}, } - Coding Agents for Generalized Task and Motion Planning ProblemsSep 2026
Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce planning effort on new instances. However, existing methods require substantial TAMP-specific engineering. We investigate whether coding agents can automate this process by synthesizing programs that generalize across instances. Given a task description and simulator access, each agent chooses how to interact with the environment while developing a program within a fixed synthesis budget. The program is then frozen and evaluated on unseen instances. We evaluate Claude Code (Opus 5) and Codex (GPT-5.6 Sol and GPT-6 Astra) on 28 simulated environments from KinDER and PDDLStream, with object counts beyond those evaluated in the original benchmark. Across all program synthesis methods, we evaluate 980 generated programs on 100 held-out instances each, 98,000 evaluation episodes in total. Overall, we find that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success (56% to 95% versus 47% for the planners, on the 16 environments where a planner is available). As object counts grow, the agents’ programs maintain higher success than the planner, using an order of magnitude less computation per instance on average. Logs show agents using interaction to calibrate physical models, test edge cases, and refine strategies. We release all code, including the full prompts given to the agents. These findings suggest that coding agents are a strong baseline for generalized TAMP.
@misc{merler2026coding, title = {Coding Agents for Generalized Task and Motion Planning Problems}, author = {Merler, Matteo and Li, Bowen and Roy, Josh and Liang, Yichao and Wang, Qianwei and Huang, Yixuan and Silver, Tom}, year = {2026}, month = sep, cv_date = {2026-09-24}, eprint = {2609.30233}, archiveprefix = {arXiv}, primaryclass = {cs.RO}, url = {https://arxiv.org/abs/2609.30233}, } - QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM AgentsSergio Hernández-Gutiérrez, Matteo Merler*, Ilze Amanda Auzina*, Joschka Strüber, Ameya Prabhu, and Matthias BethgeIn Advances in Neural Information Processing Systems (Evaluations and Datasets Track), Dec 2026To appear
LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-only rewards provide too sparse guidance, failing to inform the model about the goodness of intermediate actions. Dense supervision methods aim to solve this problem by scoring intermediate steps, from intrinsic confidence to self-distillation and embedding similarities. However, it is common practice to evaluate them by measuring the downstream performance of a training pipeline that integrates them. This is expensive, conflates supervision quality with training engineering confounders, and renders different methodological families requiring distinct training setups incomparable. As a result, dense supervision methods are rarely benchmarked on common ground. We introduce QVal, a training-free testbed for directly evaluating dense supervision signals. Given a state-action pair, QVal measures how well a method’s score is Q-aligned: whether it orders actions according to the Q-values of a strong reference-policy. This lets us compare signals before any training run and separate signal quality from other engineering choices. We instantiate QVal as QVal-v1.0, benchmarking 21 dense supervision methods across four diverse environments and seven methodological families, with over 1.2K evaluation experiments across six open-weight model backbones. We find that simple prompting baselines consistently outperform recent dense supervision methods from the literature, and that performance clusters strongly by family. These findings hold across model sizes, environments, and observation modalities. QVal is designed to be easily extensible to new environments and methods, enabling researchers to iterate on dense supervision methods before any training run.
@inproceedings{hernandez2026qval, title = {QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents}, author = {Hernández-Gutiérrez, Sergio and Merler, Matteo and Auzina, Ilze Amanda and Strüber, Joschka and Prabhu, Ameya and Bethge, Matthias}, booktitle = {Advances in Neural Information Processing Systems (Evaluations and Datasets Track)}, year = {2026}, month = dec, note = {To appear}, cv_venue = {Advances in Neural Information Processing Systems 39 (NeurIPS 2026), Evaluations and Datasets Track}, cv_date = {2026-12-06}, eprint = {2606.32034}, archiveprefix = {arXiv}, primaryclass = {cs.LG}, url = {https://arxiv.org/abs/2606.32034}, } - EMNLP 2026
Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM TeachersIn Findings of the Association for Computational Linguistics: EMNLP 2026, Oct 2026To appearVision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as policies is expensive and brittle: they must be queried at every step, do not improve from environment interaction, and can repeat systematic errors. We study how to learn a cheap autonomous policy from an online, expensive, and imperfect but informative VLM teacher. We propose SAGE (Selective Agent Guidance via Entropy), a framework that queries a VLM only when the learner is uncertain, executes the suggested action during training, and distills guidance into a lightweight Reinforcement Learning (RL) policy. Because VLM advice is not always reliable, SAGE can weight teacher-action distillation using environment-derived advantages rather than treating all suggestions as equally useful. Across sparse-reward visual reasoning and navigation tasks, SAGE learns policies that act without VLM guidance at evaluation time and improves over unguided RL in several environments, including settings where the learned policy exceeds its VLM teacher. The results show that selective guidance is most beneficial when the VLM can help the agent discover high-reward trajectories, and less useful when unguided exploration already succeeds or teacher actions do not lead to informative experience. SAGE also reduces VLM usage by prompting the teacher only on a fraction of training steps and requiring no VLM calls at deployment. Overall, our results suggest that VLMs don’t need to be used as fixed policies to be useful; they can instead act as temporary, imperfect sources of guidance whose value is tested and internalized through interaction.
@inproceedings{bonetta2026sage, title = {Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers}, author = {Bonetta, Giovanni and Merler, Matteo and Zago, Davide and Cancelliere, Rossella and Magnini, Bernardo}, booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026}, publisher = {Association for Computational Linguistics}, address = {Budapest, Hungary}, year = {2026}, month = oct, note = {To appear}, cv_venue = {Findings of the Association for Computational Linguistics: EMNLP 2026}, cv_date = {2026-10-24}, eprint = {2609.01567}, archiveprefix = {arXiv}, primaryclass = {cs.AI}, url = {https://arxiv.org/abs/2609.01567}, } - EMNLP 2026
DecSelfMask: Leveraging Unlabeled Text via Self-Relevance-Guided Masking for Decoder-Only ClassificationIn Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026), Oct 2026To appearClassification tasks require annotated data, which can often be expensive, time-consuming, or even unfeasible to collect. This is the case of the medical domain, where large datasets often have few annotated examples. To address this, we propose DecSelfMask (Decoder Self-learning by Masking), an approach to enhance decoder-only performance on classification tasks. We build on common self-learning approaches by leveraging a model to create training examples from unlabeled data to propose a novel relevance-guided masking strategy. We use relevance attribution methods to determine what portions of unannotated texts are relevant for a task. We then create self-supervised training examples by masking out those portions, training the model to reconstruct them via next-token-prediction. We hypothesize that those examples convey knowledge about the structure and semantics of unannotated data that can be useful for downstream performance. We test our approach on 136 tasks from a collection of 1.9M clinical notes from an Italian hospital. We quantify DecSelfMask’s impact on downstream tasks on 5 models of different scales and families, including a probing analysis. Experiments show consistent gains, outperforming standard supervised fine-tuning approaches (+19.9 points in Macro F1), synthetic label generation (+12.5), and continual pretraining (+6.3), as well as common baselines.
@inproceedings{ferrazzi2026decselfmask, title = {DecSelfMask: Leveraging Unlabeled Text via Self-Relevance-Guided Masking for Decoder-Only Classification}, author = {Ferrazzi, Pietro and Merler, Matteo and Bonetta, Giovanni and Lavelli, Alberto and Magnini, Bernardo}, booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)}, publisher = {Association for Computational Linguistics}, address = {Budapest, Hungary}, year = {2026}, month = oct, note = {To appear}, cv_venue = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)}, cv_date = {2026-10-24}, eprint = {2606.09466}, archiveprefix = {arXiv}, primaryclass = {cs.CL}, url = {https://arxiv.org/abs/2606.09466}, } - EMNLP 2026
ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language ModelsMatteo Merler*, Nicola Dainese*, Minttu Alakuijala, Giovanni Bonetta, Pietro Ferrazzi, Yu Tian, Bernardo Magnini, and Pekka MarttinenIn Findings of the Association for Computational Linguistics: EMNLP 2026, Oct 2026To appearIntegrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans compared to planning in natural language, with recent works extending this idea to visual domains using Vision-Language Models (VLMs). However, rigorous comparison between VLM-grounded symbolic approaches and methods that plan directly with a VLM has been hindered by a lack of common environments, evaluation protocols and model coverage. We introduce ViPlan, the first open-source benchmark for Visual Planning with symbolic predicates and VLMs. ViPlan features a series of increasingly challenging tasks in two domains: a visual variant of the classic Blocksworld planning problem and a simulated household robotics environment. We benchmark nine open-source VLM families across multiple sizes, along with selected closed models, evaluating both VLM-grounded symbolic planning and using the models directly to propose actions. We find symbolic planning to outperform direct VLM planning in Blocksworld, where accurate image grounding is crucial, whereas the opposite is true in the household robotics tasks, where commonsense knowledge and the ability to recover from errors are beneficial. Finally, we show that across most models and methods, there is no significant benefit to using Chain-of-Thought prompting, suggesting that current VLMs still struggle with visual reasoning.
@inproceedings{merler2025viplan, title = {ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models}, author = {Merler, Matteo and Dainese, Nicola and Alakuijala, Minttu and Bonetta, Giovanni and Ferrazzi, Pietro and Tian, Yu and Magnini, Bernardo and Marttinen, Pekka}, booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026}, publisher = {Association for Computational Linguistics}, address = {Budapest, Hungary}, year = {2026}, month = oct, note = {To appear}, cv_venue = {Findings of the Association for Computational Linguistics: EMNLP 2026}, cv_date = {2026-10-24}, eprint = {2505.13180}, archiveprefix = {arXiv}, primaryclass = {cs.AI}, url = {https://arxiv.org/abs/2505.13180}, }
2025
- Guiding Reinforcement Learning with Selective Vision-Language Model SupervisionMatteo Merler, Giovanni Bonetta, and Bernardo MagniniIn ECAI 2025 Workshop on AI-based Planning for Complex Real-World Applications (CAIPI’25), Oct 2025
We propose a framework that augments a model-free Reinforcement Learning (RL) agent with selective guidance from a pre-trained Vision-Language Model (VLM). Our system is designed to assist the RL agent, which starts from scratch and has no prior notion of the environment, by leveraging the VLM’s common-sense knowledge to support its decision making. Rather than relying on the VLM at every timestep, the agent monitors its own uncertainty during training and defers to the VLM only when it is unsure about which action to take. Uncertainty is measured using the entropy of the policy distribution, and guidance is triggered when this entropy exceeds a predefined threshold. To reduce computational overhead, we introduce a stochastic gating mechanism that limits the frequency of VLM queries, along with a cache that stores past VLM responses for reuse. Experiments show that our method leads to more stable learning dynamics compared to standard PPO, with reduced variance across runs. In the \textttFrozenLake environment, we observe that VLM guidance is primarily utilized during the early stages of training, gradually diminishing as the agent becomes more confident. This suggests that our selective guidance mechanism can support early exploration without hindering long-term autonomous behavior.
@inproceedings{merler2025guiding, title = {Guiding Reinforcement Learning with Selective Vision-Language Model Supervision}, author = {Merler, Matteo and Bonetta, Giovanni and Magnini, Bernardo}, booktitle = {ECAI 2025 Workshop on AI-based Planning for Complex Real-World Applications (CAIPI'25)}, editor = {Niggemann, Oliver and Biswas, Gautam and Micheli, Andrea and Heesch, René and Diedrich, Alexander and Ehrhardt, Jonas and Widulle, Niklas}, publisher = {CEUR Workshop Proceedings}, pages = {39--51}, year = {2025}, month = oct, address = {Bologna, Italy}, cv_venue = {ECAI Workshop on AI-based Planning for Complex Real-World Applications (CAIPI'25)}, cv_date = {2025-10-25}, }
2024
- Generating Code World Models with Large Language Models Guided by Monte Carlo Tree SearchIn Advances in Neural Information Processing Systems, Oct 2024
In this work we consider Code World Models, world models generated by a Large Language Model (LLM) in the form of Python code for model-based Reinforcement Learning (RL). Calling code instead of LLMs for planning has potential to be more precise, reliable, interpretable, and extremely efficient. However, writing appropriate Code World Models requires the ability to understand complex instructions, to generate exact code with non-trivial logic and to self-debug a long program with feedback from unit tests and environment trajectories. To address these challenges, we propose Generate, Improve and Fix with Monte Carlo Tree Search (GIF-MCTS), a new code generation strategy for LLMs. To test our approach in an offline RL setting, we introduce the Code World Models Benchmark (CWMB), a suite of program synthesis and planning tasks comprised of 18 diverse RL environments paired with corresponding textual descriptions and curated trajectories. GIF-MCTS surpasses all baselines on the CWMB and two other benchmarks, and we show that the Code World Models synthesized with it can be successfully used for planning, resulting in model-based RL agents with greatly improved sample efficiency and inference speed.
@inproceedings{dainese2024generating, author = {Dainese, Nicola and Merler, Matteo and Alakuijala, Minttu and Marttinen, Pekka}, booktitle = {Advances in Neural Information Processing Systems}, editor = {Globerson, A. and Mackey, L. and Belgrave, D. and Fan, A. and Paquet, U. and Tomczak, J. and Zhang, C.}, pages = {60429--60474}, publisher = {Curran Associates, Inc.}, title = {Generating Code World Models with Large Language Models Guided by Monte Carlo Tree Search}, url = {https://proceedings.neurips.cc/paper_files/paper/2024/hash/6f479ea488e0908ac8b1b37b27fd134c-Abstract-Conference.html}, volume = {37}, year = {2024}, cv_venue = {Advances in Neural Information Processing Systems 37 (NeurIPS 2024)}, cv_date = {2024-12-13}, } - In-Context Symbolic Regression: Leveraging Large Language Models for Function DiscoveryIn Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), Aug 2024
State of the art Symbolic Regression (SR) methods currently build specialized models, while the application of Large Language Models (LLMs) remains largely unexplored. In this work, we introduce the first comprehensive framework that utilizes LLMs for the task of SR.We propose In-Context Symbolic Regression (ICSR), an SR method which iteratively refines a functional form with an LLM and determines its coefficients with an external optimizer. ICSR leverages LLMs’ strong mathematical prior both to propose an initial set of possible functions given the observations and to refine them based on their errors.Our findings reveal that LLMs are able to successfully find symbolic equations that fit the given data, matching or outperforming the overall performance of the best SR baselines on four popular benchmarks, while yielding simpler equations with better out of distribution generalization.
@inproceedings{merler2024incontext, title = {In-Context Symbolic Regression: Leveraging Large Language Models for Function Discovery}, author = {Merler, Matteo and Haitsiukevich, Katsiaryna and Dainese, Nicola and Marttinen, Pekka}, editor = {Fu, Xiyan and Fleisig, Eve}, booktitle = {Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop)}, month = aug, year = {2024}, cv_venue = {Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Volume 4: Student Research Workshop (ACL 2024 SRW)}, address = {Bangkok, Thailand}, publisher = {Association for Computational Linguistics}, url = {https://aclanthology.org/2024.acl-srw.49/}, doi = {10.18653/v1/2024.acl-srw.49}, pages = {427--444}, }
* Denotes equal contribution