AIRLab · Politecnico di Milano

Alessandro Montenegro

Ph.D. Candidate in Information Technology

Reinforcement learning researcher working on the sample efficiency of policy-based methods — bridging theoretical foundations with real-world robotics and industrial applications.

Email CV Google Scholar GitHub LinkedIn
Alessandro Montenegro
01

About

I am a Ph.D. student in Information Technology at Politecnico di Milano (expected graduation: December 2026), working under the supervision of Profs. Alberto Maria Metelli, Matteo Papini, and Marco Mussi. I hold a Master's degree in Computer Science and Engineering from Politecnico di Milano (2023) and a Bachelor's degree from the University of Rome Tor Vergata (2020), both achieved with honors.

My current research focuses on Reinforcement Learning — specifically investigating the sample efficiency of policy-based methods in real-world-like settings. My work bridges theoretical foundations with practical applications, including industrial projects and robot learning.

Quick facts
Position
Ph.D. Candidate (expected Dec. 2026)
Department
DEIB (Department of Electronics, Information and Bioengineering)
Institution
Politecnico di Milano
Location
Milan, Italy
02

Research Interests

Bridging the theory-practice gap in policy-based methods of reinforcement learning — from convergence guarantees to sample-efficient, real-world deployment.

Policy Gradient Convergence
Theoretical convergence guarantees for policy gradient methods, including settings with dynamic stochasticity.
Constrained Reinforcement Learning
Policy gradient methods for constrained MDPs under continuous state-action spaces and risk-based constraints.
Sample-Efficient Trajectory Reuse
Reusing past trajectories and samples to accelerate policy gradient and proximal policy optimization training.
Robot Learning
Humanoid locomotion and foothold tracking via reinforcement learning.
03

Publications

2026
[C1]
Reusing Trajectories in Policy Gradients Enables Fast Convergence
Alessandro Montenegro, Federico Mansutti, Marco Mussi, Matteo Papini, Alberto Maria Metelli
ICML 2026 — International Conference on Machine Learning (A* Core Ranking, acceptance rate 26.6%)
Paper arXiv Code
[C2]
Mind Your Steps: A General Learning Framework for Accurate Humanoid Foothold Tracking
Alessandro Montenegro, Shihao Li, Puze Liu, Alberto Maria Metelli, Jan Peters
RSS 2026 — Robotics: Science and Systems (A* Core Ranking, acceptance rate 29.7%)
2025
[C3]
Convergence Analysis of Policy Gradient Methods with Dynamic Stochasticity
Alessandro Montenegro, Marco Mussi, Matteo Papini, Alberto Maria Metelli
ICML 2025 — International Conference on Machine Learning (A* Core Ranking, acceptance rate 26.9%)
2024
[C4]
Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning
Alessandro Montenegro, Marco Mussi, Matteo Papini, Alberto Maria Metelli
NeurIPS 2024 — Advances in Neural Information Processing Systems (A* Core Ranking, acceptance rate 25.8%)
[C5]
Learning Optimal Deterministic Policies with Stochastic Policy Gradients
Spotlight · top 3.5%
Alessandro Montenegro, Marco Mussi, Alberto Maria Metelli, Matteo Papini
ICML 2024 — International Conference on Machine Learning (A* Core Ranking, acceptance rate 27.5%)
[C6]
Best Arm Identification for Stochastic Rising Bandits
Spotlight · top 3.5%
Marco Mussi, Alessandro Montenegro, Francesco Trovò, Marcello Restelli, Alberto Maria Metelli
ICML 2024 — International Conference on Machine Learning (A* Core Ranking, acceptance rate 27.5%)
Preprints
[R1]
Reusing Samples in Proximal Policy Optimization: When and How Does It Help?
Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli
Under review, 2026
Paper arXiv Code
[R2]
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes
Alessandro Montenegro, Leonardo Cesani, Marco Mussi, Matteo Papini, Alberto Maria Metelli
Under review, 2025
Paper arXiv Code

Full list including workshop papers available in the CV.

04

Teaching

Informatica A (Computer Science)
Teaching Assistant · 10 CFU course for first-year Mathematical Engineering students · Course responsible: Prof. Giacomo Boracchi
Sep. 2024 – Aug. 2025
Exercise lecturer, written exam preparation and correction, oral exam execution · 38h · ≈150 students · Student ranking 3.3/4 (avg. university 3.1/4)
Course website ↗
05

News & Updates

July 2026
Papers accepted at ICML 2026 (Seoul) and RSS 2026 (Sydney).
December 2025
Completed a Visiting Research Fellowship at DFKI and TU Darmstadt, working with Prof. Puze Liu and Prof. Jan Peters on humanoid locomotion.
September 2025
Started a Visiting Research Fellowship at DFKI and TU Darmstadt.
2025
Paper accepted at ICML 2025.
2024
Paper accepted at NeurIPS 2024.
July 2024
Presented a poster at DLRLSS (Deep Learning and Reinforcement Learning Summer School), Toronto, Canada.
2024
Two papers accepted at ICML 2024, one as a spotlight paper (top 3.5%).
September 2023
Participated in EWRL (European Workshop on Reinforcement Learning) 2023, Brussels, Belgium.
September 2023
Started the Ph.D. in Information Technology at Politecnico di Milano.
06 · Contact

Get in touch

Feel free to reach out about collaborations, teaching, or anything else.

alessandro.montenegro@polimi.it CV Google Scholar GitHub LinkedIn
CV last updated: September 9, 2026
© 2026 Alessandro Montenegro.