Date of Submission
6-10-2026
Date of Award
6-19-2026
Institute Name (Publisher)
Indian Statistical Institute
Document Type
Master's Dissertation
Degree Name
Master of Technology
Subject Name
Computer Science
Department
Computer Vision and Pattern Recognition Unit (CVPR-Kolkata)
Supervisor
Garain, Utpal
Abstract (Summary of the Work)
Large language models have shown immense improvement in coding and math performances thanks to reinforcement learning boosted algorithms. However, its true impact on broadening the reasoning and analytical capacities of an LLM is still contended. In this dissertation, we outline the foundations of Large Language Models, and delve into Reinforcement Learning with Verifiable Rewards (RLVR). We discuss various strategies to efficiently manipulate memory during a fine tuning update. We finally perform RLVR fine-tuning techniques on different models with varied use cases and compare their performances, which corroborate the efficiency of RLVR.
Control Number
CS2423
DOI
https://dspace.isical.ac.in/items/c362b698-6e48-4f65-94e2-ca279fb81a05
DSpace Identifier
http://hdl.handle.net/10263/7768
Recommended Citation
Konnur, Rashmi, "An Empirical Study of RLVR Fine-Tuning for Mathematical Problem Solving in LLMs" (2026). Master’s Dissertations. 483.
https://digitalcommons.isical.ac.in/masters-dissertations/483