Work with me
Fixed fee, fixed scope, fixed timeline. I take two outside projects at a time and I'd rather tell you a problem isn't worth solving than bill you for finding out.
What I sell
LLM Cost Cut
Your inference bill drops at least 30%, or you don't pay. Complexity-based routing so easy queries stop hitting a frontier model, plus a two-tier semantic cache so paraphrased repeats never hit the API at all.
- Spend map: cost broken down by route, model and cache-miss reason
- Embedding + heuristic complexity router, A/B'd against your current path
- L1 in-process and L2 Redis cache with semantic similarity matching
- Dashboards for cost per request, hit rate, p95 latency and quality delta
Best if you're spending north of $8k/month on inference.
Agent Reliability Audit
Your agent works in the demo and fails for real users. I pull your production traces and tell you exactly why, ranked by how often it happens.
- Failure taxonomy with measured frequencies, not guesses
- Golden eval set built from your real traffic, wired into CI
- Ranked fix list with effort estimates
- The two cheapest high-impact fixes shipped as PRs
Best if you have an LLM feature in production and no eval suite.
RAG / Search Rescue
Users search for the exact thing and don't find it. Almost always pure vector recall with no keyword arm, so exact-match queries fall straight through.
- Judged eval set built first, so every change is measured
- Structure-aware chunking and ingestion fixes
- Hybrid BM25 + dense retrieval with fusion and metadata filters
- Latency-budgeted cross-encoder rerank
Recall@10 improves 25% on your own query set, or I keep working free.
Why me
Search and agents in production
I'm an SDE at Tejas AI (YC W25), where I own hybrid keyword + vector search on OpenSearch over a corpus indexed from around 60 ATS platforms, and the WhatsApp agent in front of it — a router model with specialist sub-agents and cross-session memory.
Inference cost, in Rust
I built an open-source inference platform that classifies query complexity via embeddings and routes accordingly, with an L1 (Moka) and L2 (Redis) semantic cache and a full Prometheus/Grafana observability stack. You can read the code.
Research to signed contracts
Co-founded AuraX as CTO — diffusion-based virtual try-on incubated at IIIT-Hyderabad's CIE, taken from research to production inference and into commercial deals with Aditya Birla Fashion Group and Mensa Brands.
How it works
- 1Twenty minutes on a call. You show me the failure, I tell you what's causing it.
- 2Within 24 hours you get a one-page scope: phases, deliverables, explicit exclusions, price, start date.
- 350% to start, 50% on delivery, net-15. Anything outside the numbered scope is a new scope.
- 4Updates every two days: shipped, next, blocked. I deliver a day early.
- 5Handover call plus a written runbook and a recording, so the people who never met me can still use it.
Tell me what's broken
One paragraph on what's failing and roughly what it costs you. If I can fix it, you'll have a scope and a price the same day. If I can't, I'll say so and point you at someone who can.
hyattherate2005@gmail.com