~/himanshu
$whoami

Work with me

Fixed fee, fixed scope, fixed timeline. I take two outside projects at a time and I'd rather tell you a problem isn't worth solving than bill you for finding out.

What I sell

LLM Cost Cut

$4,000 – $6,0002 weeks

Your inference bill drops at least 30%, or you don't pay. Complexity-based routing so easy queries stop hitting a frontier model, plus a two-tier semantic cache so paraphrased repeats never hit the API at all.

  • Spend map: cost broken down by route, model and cache-miss reason
  • Embedding + heuristic complexity router, A/B'd against your current path
  • L1 in-process and L2 Redis cache with semantic similarity matching
  • Dashboards for cost per request, hit rate, p95 latency and quality delta

Best if you're spending north of $8k/month on inference.

Agent Reliability Audit

$2,500 – $4,0005 days

Your agent works in the demo and fails for real users. I pull your production traces and tell you exactly why, ranked by how often it happens.

  • Failure taxonomy with measured frequencies, not guesses
  • Golden eval set built from your real traffic, wired into CI
  • Ranked fix list with effort estimates
  • The two cheapest high-impact fixes shipped as PRs

Best if you have an LLM feature in production and no eval suite.

RAG / Search Rescue

$6,000 – $10,0002 weeks

Users search for the exact thing and don't find it. Almost always pure vector recall with no keyword arm, so exact-match queries fall straight through.

  • Judged eval set built first, so every change is measured
  • Structure-aware chunking and ingestion fixes
  • Hybrid BM25 + dense retrieval with fusion and metadata filters
  • Latency-budgeted cross-encoder rerank

Recall@10 improves 25% on your own query set, or I keep working free.

Why me

Search and agents in production

I'm an SDE at Tejas AI (YC W25), where I own hybrid keyword + vector search on OpenSearch over a corpus indexed from around 60 ATS platforms, and the WhatsApp agent in front of it — a router model with specialist sub-agents and cross-session memory.

Inference cost, in Rust

I built an open-source inference platform that classifies query complexity via embeddings and routes accordingly, with an L1 (Moka) and L2 (Redis) semantic cache and a full Prometheus/Grafana observability stack. You can read the code.

Research to signed contracts

Co-founded AuraX as CTO — diffusion-based virtual try-on incubated at IIIT-Hyderabad's CIE, taken from research to production inference and into commercial deals with Aditya Birla Fashion Group and Mensa Brands.

How it works

  1. 1Twenty minutes on a call. You show me the failure, I tell you what's causing it.
  2. 2Within 24 hours you get a one-page scope: phases, deliverables, explicit exclusions, price, start date.
  3. 350% to start, 50% on delivery, net-15. Anything outside the numbered scope is a new scope.
  4. 4Updates every two days: shipped, next, blocked. I deliver a day early.
  5. 5Handover call plus a written runbook and a recording, so the people who never met me can still use it.

Tell me what's broken

One paragraph on what's failing and roughly what it costs you. If I can fix it, you'll have a scope and a price the same day. If I can't, I'll say so and point you at someone who can.

hyattherate2005@gmail.com