MS in Data Science at UC San Diego · graduating December 2026
Machine Learning Engineering Intern at Infoblox (Summer 2026)
LinkedIn · Medium · contextjetai.com/nishchay · emailfornishchay@gmail.com
I build production AI systems and the unglamorous infra that keeps them upright. I get nerd-sniped by anything at the intersection of evals, agent orchestration, and inference cost.
Languages Python, SQL, TypeScript, C++, Bash ML & modeling PyTorch, TensorFlow, scikit-learn, XGBoost, CatBoost, Prophet, MLflow LLM tooling LangChain, LangGraph, LlamaIndex, DSPy, MCP, Hugging Face, OpenAI, Anthropic Vector & graph Pinecone, Milvus, Weaviate, FAISS, Neo4j Infra & delivery Docker, Kubernetes, FastAPI, GitHub Actions, Azure, AWS, GCP, Databricks, Supabase Frontend & data viz React, Next.js, Streamlit, Plotly, D3.js, Dash
- Shipping fixes and features into the AI tooling I actually use day to day. 30 merged PRs across 20 orgs so far — mostly provider integrations and correctness fixes found by differential-fuzzing hand-rolled parsers against the stdlib. Highlights below.
- Building and maintaining my own tools under ContextJet-ai:
awesome-llm-observability — 50+ curated observability tools plus 26 installable agent skills — and mcpvitals, a one-command health check for MCP servers (on PyPI): health score, token cost, tool-confusion and migration readiness.
- Back at UC San Diego for the last stretch of the MS in Data Science, graduating December 2026 and looking for ML / AI engineering roles.
- Just wrapped a summer at Infoblox as an ML engineering intern — a probability-to-renew model for a flagship product line (0.89 AUC) feeding renewal-risk prioritisation for sales, customer health scoring across every active account, and research into the agentic AI roadmap for their DNS security product.
- Reading the LangGraph internals, whatever new agent paper is going viral that week, and the older systems books that age well (Designing Data-Intensive Applications stays open on my desk).
The merges I'd show first:
Apple MLX —
top_psmall enough to round past float precision masked every token, so the sampler returned noise instead of the most likely one. Worst at bfloat16, which is what on-device inference actually runs.
Comet Opik — four merges: Groq and Cerebras integrations, the Ollama SDK integration, and a logging fix that was swallowing every stream diagnostic
dottxt outlines — two type-system fixes found by differential-fuzzing the hand-rolled regexes against
ipaddress— leading-zero octets inipv4and the same bug in IPv4-mapped IPv6 — plus a mypy hook fix for the style job
NVIDIA garak — native Anthropic generator for the LLM vulnerability scanner
dify — storage-layer
@overriderefactor in the most-starred open-source LLM app platform
Stanford DSPy —
Document.format()was emitting an invalid source type for PDFs
Google A2A — a silently-ignored
queue_managerin the v2 request handler, with the warning the project's own test harness needed
promptfoo — three provider integrations: NVIDIA NIM, Fireworks AI, and Moonshot Kimi
Arize openinference — Cohere, Ollama and Together AI instrumentors
Mistral AI — from_modeldeprecation fix in the official tokenizer library
AWS agentcore-cli — zip-stage config regression fix
SQLMesh — invalidating a nonexistent environment failed silently instead of erroring
Also merged into
Weaviate,
Deepgram,
Voyage AI,
Cartesia,
Braintrust,
Baseten Truss,
Mirascope and ogx.
Still open in
openllmetry,
aider, the
MCP TypeScript SDK,
vLLM,
Apple coremltools,
pinecone and a few more.
| Project | What it does | Impact |
|---|---|---|
| Knowledge GraphRAG Platform | Entity-linked graph over docs for import/export compliance. LangGraph, vector DB, Salesforce. | +87% answer precision, −45% research time. Auditable citations. |
| Multimodal Synthetic Market Surveys (C5i.ai) | Real-time respondent synthesis for CPG and marketing studies. LangChain, Azure OpenAI, multi-agent, Apify. | ~90% accuracy vs live benchmarks. $300K+ in attributable revenue. |
| AI Sales Development Representative (Wall Street client) | Prospecting, enrichment, personalization, outreach, reply handling for a PE / hedge-fund / family-office target list. LangChain agents, Pydantic workflows, React. | +35% qualified meetings, −60% manual prospecting, 200–300 leads/week. |
| LLM Virtual Try-On Assistant (apparel client) | Diffusion-based try-on (StableVITON) + OpenAI image + LangGraph + MediaPipe + Pinecone RAG over catalog. | Time-on-page +25%, CTR +18%. |
| Predictive Maintenance + RAG (Industry 4.0) | Vibration/temperature anomaly detection + RAG + forecasting for conveyor planners. scikit-learn, Prophet, LangGraph, Databricks. | Unplanned maintenance −15%, planning cycle −30%. |
- Evals are the only thing that scales engineering judgment. Most teams write the eval after deciding the model is good, which is backwards.
- Agent frameworks are mostly thin glue. Read the source before you adopt one.
- The best LangChain users I know also use less of LangChain over time.
- The cheapest performance win is almost always a smaller, better prompt. The second cheapest is caching. Quantization is rarely the answer people think it is.
- REAL MADRID and CR7.
When I get a weekend and a problem statement, I build things like StepWise (AWS Breaking Barriers 2024, digital inclusion copilot), DocuGuard AI (HackAI Dell/NVIDIA 2024, enterprise document risk), and NeuroForecast AI (UCSD SMASH NSF HDR 2026, OOD-robust neural forecasting). Constraints make for sharper systems.
- Revolutionizing Market Surveys through Generative AI for Efficient Data Synthesis, keynote at the Machine Learning Developers Summit 2024, published in Lattice Journal (AoDS), Vol. 5 Issue 1.
- Occasional notes on Medium.


