Author (Researcher Name)

Date of Submission

6-10-2026

Date of Award

6-19-2026

Institute Name (Publisher)

Indian Statistical Institute

Document Type

Master's Dissertation

Degree Name

Master of Technology

Subject Name

Computer Science

Department

Computer Vision and Pattern Recognition Unit (CVPR-Kolkata)

Supervisor

Garain, Utpal

Abstract (Summary of the Work)

Large language models have shown immense improvement in coding and math performances thanks to reinforcement learning boosted algorithms. However, its true impact on broadening the reasoning and analytical capacities of an LLM is still contended. In this dissertation, we outline the foundations of Large Language Models, and delve into Reinforcement Learning with Verifiable Rewards (RLVR). We discuss various strategies to efficiently manipulate memory during a fine tuning update. We finally perform RLVR fine-tuning techniques on different models with varied use cases and compare their performances, which corroborate the efficiency of RLVR.

Control Number

CS2423

DOI

https://dspace.isical.ac.in/items/c362b698-6e48-4f65-94e2-ca279fb81a05

DSpace Identifier

http://hdl.handle.net/10263/7768

Share

COinS