NOC / Deep Learning Intern · Sterlite Technologies
Prototype
You are the human in RLHF — pick the better assistant response, round 1/5
?
- Reward-model accuracy
- Projected MTTR
- Preference pairs
- NOC alone 🎫🎫🎫🎫🎫🎫🎫🎫
- NOC + LLM assistant 🎫🎫🎫🎫🎫🎫🎫🎫
Stylized miniature of the real pipeline — SFT over 10,000 prompt-response pairs, a binary reward classifier trained on 30,000 preference pairs, then policy optimization: the deployed assistant cut real mean time to resolution by ~35% for 100+ enterprise users. Run the shift comparison before training the reward model and the assistant barely helps.
Highlights
- Enterprise LLM Integration: Successfully accelerated customer management workflows by integrating an LLM-powered self-service conversational assistant into an enterprise OSS/NMS Network Management System serving 100+ active enterprise users.
- Pretraining & SFT: Conducted generative pretraining scaling tests over 200 Million Kaggle text samples and executed Supervised Fine-Tuning (SFT) over 10,000 high-quality human and AI prompt-response pairs.
- Reward Modeling & RLHF: Engineered an alignment reward pipeline utilizing a binary classifier trained on 30,000 preference learning pairs, assigning scores based on accurate user configuration criteria.
- Reinforcement Learning: Applied policy optimization techniques to train response generations based on preference scores, achieving target accuracy and slashing mean time to resolution (MTTR) by ~35%.
- Automation Engineering: Automated massive multi-GPU data formatting pipelines via native Bash/Python automation wrappers, increasing core processing runtimes by 35%.
- Infrastructure Support: Worked alongside network operations engineers to diagnose and resolve 20+ real-time critical routing, switched fabric, DNS, and DHCP incidents, bringing resolution speed down from ~2 hours to under 45 minutes while co-authoring internal runbooks.
Technologies
- LLM Integration
- SFT
- RLHF
- Reinforcement Learning
- Bash / Python
- DNS / DHCP