
research note
ExpRL — Exploratory RL for LLM Mid-Training
This paper addresses the challenge of improving large language model (LLM) reasoning capabilities through reinforcement learning (RL) when sparse reward signals are insufficient due to limited base…










