Research
cool research stuff
Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs SFT Finetuned Models
Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit, Kevin Zhu, Aishwarya Balwani
Investigating how reinforcement learning and supervised fine-tuning shape internal representations for mathematical reasoning.