14000 ft. Tawang, India
Fourteen thousand feet in Tawang, Arunachal — mist on the hills, a river underfoot, and no plans except go. Trails beat parties, every time.


✻ samunder singh
Bengaluru, Karnataka, India
Design Engineer at C-DAC Bangalore
GPU Design Engineer / AI/ML Engineer / Open Source Contributor
Building LLMs for the edge, Vision AI, and the software stack behind custom GPU hardware.
I'm an AI/ML Design Engineer working at the intersection of LLMs, distributed systems, GPU computing, and AI infrastructure. Currently at C-DAC Bengaluru, I build 1B–7B small language models with a focus on distributed training, inference optimization, quantization, and deployment. I also work on the software stack for custom GPU hardware — OpenCL kernels, RISC-V, POCL, compiler/runtime development — and enabling PyTorch-like AI workloads on custom accelerators.
Previously at Lincode I built and deployed production ML/CV systems using PyTorch, YOLO, DETR, CLIP, TensorRT, torch.compile, and distributed training. I'm interested in LLM systems, GPU programming, compilers, AI accelerators, and hardware–software co-design. I also contribute to open source, including Ivy at Unify.
B.Tech in Electrical, Electronics and Communications Engineering · Tezpur University, Assam, India · 2020 - 2024
July 2026 - Present
Bengaluru, India
September 2024 - July 2026
Bengaluru, India
May 2024 - September 2024
United States (Remote)
February 2024 - May 2024
Remote
Deep Learning
A high-performance deep learning framework on PyTorch with custom shader-based kernels through Slang, for efficient tensor operations on GPU.
Slang PyTorch Python CUDA
Computer Vision
Real-time Indian Sign Language recognition with Mediapipe, CNN, and LSTM at 90%+ accuracy. Won an award at a Hasgeek hackathon.
Python MediaPipe TensorFlow Flask
Computer Vision / NLP / LLM
FastAPI video analysis platform using ML for figure detection, OCR, and speech-to-text, with Google OAuth and Vectara.
FastAPI Celery Google OAuth VectorDB
NLP / LLM
Chatbot Arena for benchmarking LLMs in the wild — enter a prompt and compare two models side by side.
Docker GCP Streamlit LLMs
Computer Vision / LLM
LangChain agent that talks to Ultralytics YOLO so you can train, validate, export, and deploy detection models in natural language.
YOLO LangChain Python OpenAI
Computer Vision
Translated real-time sign gestures into text using holistic landmark tracking with MediaPipe and TensorFlow.
Python MediaPipe TensorFlow Flask
NLP / LLM
LLM chatbot that fetches live weather and search results with LangChain and OpenAI.
LangChain OpenAI Streamlit RAG
Computer Vision
Custom object detection on an aquarium dataset with visual logging and monitoring.
YOLOv7 PyTorch W&B OpenCV
Samunder has done great work during his time as a top contributor working on Ivy. His enthusiasm was infectious to those around him, and he learned a lot during the role, picking up new skills very quickly, with a drive to get things done. I would recommend him for any role involving Machine Learning infrastructure, and especially those using a lot of Python!
Daniel Lenton — CEO at Unify | YC W23 · May 6, 2024
Not really a party person. Give me a game, a trail, a tent — or a hackathon.
Fourteen thousand feet in Tawang, Arunachal — mist on the hills, a river underfoot, and no plans except go. Trails beat parties, every time.


And yes — we won The Fifth Elephant hackathon. Show up, ship, take the trophy.


singhsamunder270@gmail.com · 8741804051