Skip to main content

Umarfarook Gurramkonda · founding ML engineer · HypeOn AI, Bengaluru

I measure what I ship.

I build multi-agent LLM systems and natural-language interfaces over data. I ship each one behind an eval harness and a hard cost cap, and the repos are public. Rerun any number on this page.

$ mail umarfarook0yt@gmail.com

available now · 15-day notice · remote or contract · I work US, EU or AU hours

01 / evidence

Seven numbers, all rerunnable.

14/15

exact-match field extraction from raw quote emails · 6 fields, n = 15

fig. 1.1 · Cargo Concierge ablation

0.782

YOLOv8n mAP@0.5 · 6,000-image sample, 25 epochs

fig. 1.2 · License Plate Privacy Blurring

2,680 ms

mean extraction latency · Gemini Flash, full instructions

fig. 1.3 · Cargo Concierge ablation

+33 pts

accuracy the rules block is worth · 15-item ablation

Cargo ablation

82

offline tests, no API key needed

TrustBench

100 MB

mandatory query cost cap

mcp-bigquery-evals

~4.1 ms

detector inference on a T4 · 3.0M params

License Plate Privacy Blurring

03 / method

Same four stages on every project.

[01]orchestrateI route each request through extraction, rate lookup, ranking and drafting, and stream progress while they run. One quote fans out to six model calls: extraction, up to three rationales, a recommendation and the draft email. Extraction averages 2,680 ms in the 15-item ablation. The end-to-end quote is not timed.proof · Cargo Concierge
[02]groundI put the vector store behind a Protocol so the backend is a one-file swap, and the retrieval metrics (Recall@K, MRR, nDCG@10) are a component with a CLI rather than a script run once. The harness has not been pointed at a benchmark yet. Below a confidence floor the answer comes back as “I don’t know” plus the closest passage. Citations are required or the answer is rejected.proof · RAG Document QA · v0.0.1
[03]constrainI dry-run every BigQuery call and refuse anything that would scan past 100 MB. Seven read-only tools, seven stable error codes an agent can switch on. There is no write path.proof · mcp-bigquery-evals
[04]measureI score every answer on eight metrics, five judged and three deterministic. When a slice regresses I run McNemar’s exact test before I call it real. The Cohen’s kappa function for judge-versus-human calibration is written and tested, but I have not labelled a set to run it against. 82 tests run offline with no API key.proof · TrustBench

04 / checkpoints

Three roles since 2024.

  1. Oct 2025 → now

    Founding ML Engineer

    HypeOn AI

    I own the LLM orchestration service: routing, intent and composition sub-agents streaming over SSE, plus NL-to-SQL over BigQuery that dry-runs every query before it costs anything. Claude Haiku is primary, Gemini is the fallback.

    LangChain · FastAPI · BigQuery · Cloud Run · Claude · Gemini

  2. Oct 2024 → Sep 2025

    Freelance ML / AI Engineer

    Independent

    I built an inventory system for a retail client. I extracted invoices with an LLM and forecast stock levels with scikit-learn, then put both behind one dashboard.

    Python · scikit-learn · OpenAI · SQL

  3. Jun → Sep 2024

    Backend Developer Intern

    Synclovis Systems

    I built REST services in Node, Express and MySQL for an event-management app. I also added LangChain and FAISS retrieval to an internal LLM health assistant, with guardrails that refuse out-of-scope questions.

    Node · Express · MySQL · LangChain · FAISS

education

B.Tech · Computer Science · 2024

K.S.R.M College of Engineering · 8.14 / 10

certifications

  • Oracle OCI Data Science Professional · 2025
  • Oracle OCI AI Foundations Associate · 2025
  • Azure AI Fundamentals · AWS Cloud Foundations

05 / working set

Everything here is in something I shipped.

models + agents

backend + data

delivery

interface

06 / about

Umarfarook Gurramkonda
fig. 6.1 · the source image

Most of my time goes to the layer around the model.

I decide when a model earns its place, then check the output still holds under real traffic. Most of the work is evaluation and cost discipline. Neither is fun. I want applied ML and LLM engineering roles at early-stage companies, where that judgment counts as much as model choice.

Umarfarook.

07 / contact

Write to me.

$ mail umarfarook0yt@gmail.com

I read my own inbox. I reply inside a day, most days.

available now · 15-day notice · remote or contract · I work US, EU or AU hours from Bengaluru

Resume ↗