← Haadhi Irfan

NOC / Deep Learning Intern · Sterlite Technologies

Gurugram, India · May 2023 – July 2023

Prototype

You are the human in RLHF — pick the better assistant response, round 1/5 ?

Stylized miniature of the real pipeline — SFT over 10,000 prompt-response pairs, a binary reward classifier trained on 30,000 preference pairs, then policy optimization: the deployed assistant cut real mean time to resolution by ~35% for 100+ enterprise users. Run the shift comparison before training the reward model and the assistant barely helps.

Highlights

Technologies