IIT Bombay · CSE · BharatGen · AI Research

Ajay Rangoji

LLMs, Generative AI, Deep Learning, Indic NLP, Reinforcement Learning

Technical specialist in Large Language Models, Generative AI, and Deep Learning with a rigorous mathematical foundation across linear algebra, probability, and calculus. Researcher at BharatGen — IIT Bombay's flagship Indic AI initiative — under Prof. Ganesh Ramakrishnan, building diagnostic benchmarks and reward-shaping pipelines for LLM reasoning in Hindi, Telugu, Marathi, and Bengali.

Core Specialisations

Large Language Models

Architecture internals, reasoning diagnostics, failure taxonomy, RLHF/GRPO fine-tuning, prompt engineering at scale.

Generative AI

Text generation pipelines, instruction-tuned models, retrieval-augmented generation, agentic LLM workflows.

Deep Learning

Transformer architectures, attention mechanisms, LoRA/QLoRA adaptation, model evaluation and benchmarking.

Reinforcement Learning

Policy optimisation, reward shaping, GRPO/PPO for language models, TRL library, multi-step reasoning incentivisation.

Indic NLP · BharatGen

Low-resource multilingual benchmarking, cross-lingual reasoning evaluation, native fluency in Hindi and Telugu.

Mathematical AI

Linear algebra, probability theory, statistical inference, optimisation, calculus — applied to ML model design.

Keyword Index

BharatGenLarge Language ModelsGenerative AILLM Reasoning & Failure AnalysisGRPOReward ShapingDeep LearningTransformer ArchitecturesRetrieval-Augmented GenerationReinforcement LearningIndic NLPChain-of-Thought PromptingInstruction Fine-TuningGraph Neural NetworksLinear AlgebraProbability TheoryCalculus & OptimisationStatistical InferencePyTorchHuggingFaceTRLPythonC++CUDAParam Cluster

Research

M.Tech Thesis — Multi-Teacher On-Policy Distillation for Verifiable and Unverifiable Domains

Guide: Prof. Ganesh Ramakrishnan

RL-trained domain experts, shared SFT checkpoint, per-prompt routing, token-level reverse-KL distillation, exposure-bias elimination, multi-round iterative distillation, verifiable + unverifiable domain unification

R&D Report — Empirical Study of LLM Reasoning Failures on English Benchmarks

Foundation work for the M.Tech thesis

4-model evaluation, 199-question benchmark, 10-category failure taxonomy, premise-order sensitivity, multi-step error propagation, GRPO reward shaping

Seminar — Reasoning in LLMs

Guide: Prof. Ganesh Ramakrishnan

Chain-of-Thought, Tree-of-Thoughts, recursive decomposition, RL-based reasoning, inference-time vs. training-time trade-offs, exploration-efficiency analysis

Projects

Multimodal GraphRAG for Long-Document QA

VLM captioning, knowledge graph construction, entity resolution, 2-hop traversal, community detection, MMLongBench-Doc

Reasoning Failure Diagnosis in Indic-Language LLMs

16-category failure taxonomy, 17B Indic LLM, MGSM, Hindi/Telugu/Bengali

Tapestry: AI Alliance

Federated fine-tuning, Llama 3.1 8B, multi-GPU multi-node training, WVS + IVM evaluation pipeline

Agentic Research Assistant with Multi-Source Synthesis

LangGraph state machine, autonomous agent nodes, citation-backed synthesis

MNIST Compression Robustness Analysis

Pruning, distillation, INT8 quantization, LeNet-5

Paged KV-Cache Scheduler for LLM Inference

Continuous batching, paged KV cache, GPU scheduling

CivicEase – Civic Services Aggregation Platform

Flask, SQLAlchemy, Bcrypt authentication, RSS ingestion pipeline

Multi-Threaded Key-Value Store with LRU Cache & Load Benchmarking

C++, PostgreSQL, LRU cache, mutex-guarded connection pooling

Skills & Technologies

LLMs & Gen AI

Reasoning analysis & benchmarking, instruction fine-tuning (LoRA/QLoRA), GRPO/PPO reward shaping, retrieval-augmented generation, Chain/Tree/Buffer of Thought, prompt engineering & evaluation

Deep Learning

Transformer & attention mechanisms, sequence-to-sequence models, graph neural networks, convolutional & recurrent nets, model compression & quantisation, transfer learning

ML & RL

Supervised & unsupervised learning, policy-gradient reinforcement learning, Bayesian methods, ensemble & boosting methods, evaluation metrics, experiment tracking with Weights & Biases

Mathematics

Linear algebra & matrix decompositions, probability theory & stochastic processes, statistical inference & hypothesis testing, multivariate calculus & gradients, convex optimisation, information theory

Frameworks & Libraries

PyTorch, TensorFlow, HuggingFace Transformers, TRL, PEFT, LangChain, LlamaIndex, scikit-learn, NumPy, Pandas, NLTK, spaCy, IndicNLP, MLflow, Verl, PrimeRL, Slakshana

Programming & Systems

Python, C, C++, Java, JavaScript, HTML, CSS, SQL, Bash, CUDA/GPU programming, Linux, Git, Docker, Kafka, PostgreSQL, MySQL, LaTeX, Param Supercomputer, OS, Networks, COA, DBMS, DSA

Education

M.Tech — Computer Science & Engineering

Indian Institute of Technology Bombay · In Progress

MTP under Prof. Ganesh Ramakrishnan · BharatGen Project · Param Compute Cluster Access · specialisation in LLMs, Generative AI, and Indic NLP

B.Tech — Computer Science & Engineering

R.V.R. & J.C. College of Engineering, Guntur

Strong foundation in algorithms, mathematics, and systems programming · 9.02/10 CGPA

Higher Secondary — Science (PCM)

Sri Chaitanya Junior College

~98%

Secondary Education (SSC)

Sri Satya Sai Vidya Vihar

10/10 CGPA

Grounded In

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Wei et al., 2022

Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Yao et al., 2023

Buffer of Thoughts: Thought-Augmented Reasoning with LLMs, Yang et al., 2024

Algorithm of Thoughts: Enhancing LLM Exploration via Algorithmic Reasoning, Sel et al., 2023

Multi-Step Action Planning for Complex Reasoning in LLMs

DeepSeek-R1: Incentivising Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek AI, 2025

Premise Order Matters in Question Answering for Large Language Models, Chen et al., 2024

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes, Agarwal et al., 2023

MiniLLM: Knowledge Distillation of Large Language Models, Gu et al., 2023

Reinforced Multi-Teacher Selection for Knowledge Distillation, Yuan et al., AAAI 2021

Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs, Jin et al., ICLR 2026

Achievements

GATE Qualified

GATE CS 2025 — All India Rank 289 (~99.99 percentile), cleared in B.Tech 4th year. GATE CS 2024 — All India Rank 606, cleared in 3rd year.

Researcher, BharatGen

Selected to contribute to IIT Bombay's national flagship Indic-language AI initiative, under Prof. Ganesh Ramakrishnan, with AWS cluster GPU access.

Best Student Award

Top-1 finish — Code Craft

CodeWiz, organized by ACM.

Participated in ICPC

International Collegiate Programming Contest — an international coding olympiad.

Positions of Responsibility

Reasoning Team Member — Machine Learning Lab, IIT Bombay

Fresher guidance, core ML fundamentals, LLM training phases, lab research showcase

Teaching Assistant — CS 769: Optimization in Machine Learning

Research-paper-based projects, skill-based interviews, exam invigilation

Teaching Assistant — CS 207: Discrete Structures

Doubt-clearing sessions, assignment grading, evaluation support, Moodle management

Teaching Assistant — CS 228: Logic for Computer Science

Exam leadership, grading, invigilation, 200+ student mentoring

Student Companion — Institute Student Companion Programme (ISCP)

Rigorous selection process, fresher support, academic & socio-cultural guidance, holistic well-being

Academic Coursework

M.Tech — IIT Bombay

Advanced Machine Learning, Deep Learning & Neural Networks, Natural Language Processing, Foundations of Reinforcement Learning, Graph Learning & Network Analysis, Advanced Algorithms & Complexity, Computer Vision, Research Methodology

B.Tech Foundation

Data Structures & Algorithms, Object-Oriented Programming, Database Management Systems, Computer Networks & Security, Operating Systems, Software Engineering, Theory of Computation, Probability & Statistics