# Himanshu — full site text Every page on https://himanshuat.com as Markdown, concatenated. Each section keeps the canonical URL it came from, so anything quoted here can be attributed. Any single page is also available on its own by appending `.md` to its URL. ======================================================================== # Himanshu I co-founded **AuraX**, a try-on startup whose diffusion models outperformed Google's and Stable Diffusion's in our benchmarks. Before that I built **SynthAiLabs**, a 500-student learning community that juniors restarted after I left. Most recently I built search and agents at **Tejas AI (YC W25)**. I left in August 2026 and I'm **looking for what's next**. ## Availability Left Tejas AI (YC W25) in August 2026. Open to full-time roles in AI infrastructure, search and quantitative research; contract engagements; and research collaborations. Book 15 minutes: https://cal.com/himanshuat/makecustom12?duration=15 — or email hyattherate2005@gmail.com ## Sections - [About](https://himanshuat.com/about) — Markdown: https://himanshuat.com/about.md - [Experience](https://himanshuat.com/experience) — Markdown: https://himanshuat.com/experience.md - [Projects](https://himanshuat.com/projects) — Markdown: https://himanshuat.com/projects.md - [Research](https://himanshuat.com/research) — Markdown: https://himanshuat.com/research.md - [Writing](https://himanshuat.com/blogs) — Markdown: https://himanshuat.com/blogs.md - [Books](https://himanshuat.com/books) — Markdown: https://himanshuat.com/books.md - [Skills](https://himanshuat.com/skills) — Markdown: https://himanshuat.com/skills.md - [Changelog](https://himanshuat.com/changelog) — Markdown: https://himanshuat.com/changelog.md - [Work with me](https://himanshuat.com/hire) — Markdown: https://himanshuat.com/hire.md --- Source: https://himanshuat.com/ ======================================================================== # About Himanshu I co-founded **AuraX**, a try-on startup whose diffusion models outperformed Google's and Stable Diffusion's in our benchmarks. Before that I built **SynthAiLabs**, a 500-student learning community that juniors restarted after I left. Most recently I built search and agents at **Tejas AI (YC W25)**. I left in August 2026 and I'm **looking for what's next**. ## 01 · Built a community that outlived me. *2023 – 2024* In my 2nd year at NIT Kurukshetra, I started **SynthAiLabs (OpenEdu)** — part startup, part movement. I hand-selected 500 students from 2,000 and built an accelerated learning community. Students were completing 6 months of curriculum in just 2. The revenue model was B2B: companies could distribute coupons to students, and I'd sell outsourced learning materials and software to companies training their employees. But funds were always the bottleneck. What happened after I stopped was the real validation — **juniors restarted and continued it on their own.** It had given them something worth continuing. ## 02 · Beat the benchmarks. Ran out of runway. *2024 – 2025* I co-founded **AuraX**, a fashion-tech startup incubated at **IIIT-Hyderabad's CIE**. We researched **diffusion models for virtual try-on** and built a pipeline that, in our quality benchmarks, outperformed Google's and Stable Diffusion's published results — along with the Indian competitors we tested against. We signed a deal with **Aditya Birla Fashion Group** and several brands under **Mensa Brands**. Real revenue. Real contracts. But the space moved fast. Larger players with deeper pockets entered. We ran out of runway. I made the call to close AuraX — not because the tech failed, but because the timing wasn't on our side. ## 03 · Pause. Reset. Rebuild. *early 2026* After closing AuraX, I took intentional time off. I studied finance to understand the business mistakes I'd made — and started building on [MQL5](https://www.mql5.com/en/users/himanshuat) along the way. I focused on health — something I'd ignored for two years of building. > Sometimes, the most productive thing that you can do is to step outside and do nothing... relax and enjoy nature. > > — Melanie Charlene ## 04 · Shipped agents at scale. Then left. *2026* For five months I was an SDE at **Tejas AI (YC W25)** — WhatsApp-based AI agents with cross-session memory, a scraping fleet indexing roughly 60 ATS platforms, and OpenSearch-backed hybrid retrieval doing the matching. The lesson was that none of the hard parts were model problems. They were memory, deterministic routing, and knowing when a user's requirement is a filter rather than a similarity score. I left in **August 2026**. ## 05 · Open to what's next. *Now* I'm looking for the next thing, and I'm deliberately open about the shape it takes: a **full-time role** in AI infrastructure, search, or quantitative research; **contract work** on inference cost, agent reliability, or retrieval that misses; or a **research collaboration** — I'm currently working on measurement studies in systems and finance. If any of that is a thing you need, the fastest route is [email](mailto:hyattherate2005@gmail.com). The productized versions of the contract work are on the [hire page](/hire). --- Source: https://himanshuat.com/about ======================================================================== # Experience Every role, newest first. Append `.md` to any detail URL for its Markdown. ## Software Development Engineer (SDE) — Tejas AI (YC W25) *Mar 2026 – Aug 2026 · SF, California, USA* Tejas AI builds Elip, an AI career coach that lives where job seekers already are: WhatsApp, with a web app, a Chrome extension and a mobile app around it. The pitch is that looking for work is mostly unglamorous logistics, and most of that can be handled by an agent that remembers you between conversations. Detail: https://himanshuat.com/experience/tejas-ai-yc-w25-software-development-engineer-sde.md ## AI Researcher & CTO — AuraX *Dec 2024 – Dec 2025 · Hyderabad, India* AuraX was a fashion-tech company incubated at IIIT-Hyderabad's CIE. Fashion e-commerce needs on-model photography for every garment, in every size, on every body type, and a photoshoot per SKU does not scale. We generated that imagery with diffusion models instead. Detail: https://himanshuat.com/experience/aurax-ai-researcher-cto.md ## Founder — SynthAiLabs (formerly OpenEdu) *Dec 2023 – Feb 2025 · Remote* SynthAiLabs, which started as OpenEdu, was an attempt to compress how long it takes a student to become employable. Part edtech product, part community, aimed at Indian engineering students who had the ability but not the ramp. Detail: https://himanshuat.com/experience/synthailabs-formerly-openedu-founder.md ## AI Dev — Tublian *July 2024 – Dec 2024 · Columbus, Ohio* Tublian is a developer platform, with TublianOS as the product and an open-source collaboration programme around it. I came back to it in a second, more engineering-focused stint. Detail: https://himanshuat.com/experience/tublian-ai-dev.md ## QM Researcher — Stealth *March 2024 – June 2024 · Kyoto, Japan* A stealth research group in Kyoto working at the intersection of quantum computing and machine learning. The specifics are under NDA, so this entry stays deliberately thin. Detail: https://himanshuat.com/experience/stealth-qm-researcher.md ## Open Source Collaborator | Project Mentor — Tublian *Dec. 2023 – April 2024 · Columbus, Ohio* Tublian's open-source programme pairs developers with real projects and with each other. I joined as a contributor and ended up mentoring. Detail: https://himanshuat.com/experience/tublian-open-source-collaborator-project-mentor.md ## Project Manager — Excelerate *April 2023 – Sept 2023 · Remote* Excelerate runs global education and experience programmes for students. I led delivery on one of their events. Detail: https://himanshuat.com/experience/excelerate-project-manager.md --- Source: https://himanshuat.com/experience ======================================================================== # Research Published and ongoing work. Append `.md` to any detail URL for its Markdown. ## It's the Fault Policy, Not the Page Size: Why mmap Loses to read() for Out-of-Core Scans on Apple Silicon *ongoing · Systems* Re-measuring mmap vs read() on Apple Silicon: a working-set crossover, attributed to OS fault-population policy rather than page size. Detail: https://himanshuat.com/research/mmap-vs-read-apple-silicon.md ## Conflict-Aware Adapter Composition (C-AAC) *completed · Computer Vision* Outperforms base Flux, Imagen 4, and ChatGPT in skin quality and aesthetic appeal. Detail: https://himanshuat.com/research/c-aac.md ## Flux-VTON+: Hybrid Flux Inpainting and Multi-LoRA Expert Fusion *completed · Generative AI* State-of-the-art VTON pipeline synergizing Flux Fill and Redux with Multi-LoRA Fusion, achieving an 85% success rate on complex garments. Detail: https://himanshuat.com/research/renu-virtual-try-on.md ## MoE: Brain-Inspired LLM Architecture *ongoing · Artificial Intelligence* A whitepaper advancing MoE as a computational analogue to the brain's functional specialization. Detail: https://himanshuat.com/research/moe-brain.md ## High-Dimensional Lattice-Based Quantum Encryption (LWE) *stopped · Cryptography* A PQC scheme using the LWE problem with detailed mathematical formulations. Detail: https://himanshuat.com/research/lattice-based-quantum-encryption.md ## GPT-Neo: Transformer Implementation *completed · Natural Language Processing* Implementing 'Attention is All You Need' from scratch. Detail: https://himanshuat.com/research/gpt-neo.md ## Custom Neural Network for MNIST *completed · Machine Learning* A custom-built Neural Net for the MNIST digit recognizer dataset. Detail: https://himanshuat.com/research/custom-n-n-mnist.md --- Source: https://himanshuat.com/research ======================================================================== # Projects ## CCProxy *ongoing · AI Tools · Mar 2026 -- ongoing* - A universal proxy that lets Claude Code talk to any AI provider: OpenAI, Gemini, or local Ollama models. Stack: AI Tools, TypeScript Site: https://ccproxy.himanshuat.com/ Source: https://github.com/Himasnhu-AT/ccproxy ## OpenCrow *ongoing · AI Platform · Jan 2026 -- ongoing* - Open-source AI agent platform for embedding intelligent customer support assistants into web applications. - Features an embeddable React chat widget with dual-layer tooling: server-side tools via OpenAPI specs and client-side tools executed in the browser. - Dynamic tool merging lets the LLM pick between server and client actions per request. - Includes a management dashboard for agents, products, and integrations. Inspired by usecrow.ai. Stack: React, TypeScript, Node.js, OpenAPI, LLM Site: https://opencrow.himanshuat.com Source: https://github.com/Himasnhu-AT/opencrow ## AI Inference Platform *completed · Infrastructure · Dec 2025 -- Jan 2026* - High-performance AI inference platform with intelligent model routing — classifies query complexity via embeddings and heuristics, routing simple queries to fast/cheap models and complex ones to GPT-4 class models. - Multi-level caching: L1 in-memory (Moka) + L2 distributed (Redis) with semantic similarity matching via cosine similarity on embeddings. - Full observability stack with Prometheus metrics and pre-configured Grafana dashboards for inference latency, cache hit rates, and model health. - Fully async Rust architecture with horizontal scaling via Docker and Kubernetes. Stack: Rust, Axum, Tokio, Redis, Prometheus, Grafana, Docker Site: https://infer.himanshuat.com ## TuneSLM *ongoing · AI Tools · Jan 2026 -- ongoing* - Drop your dataset and answer a few questions — get an ipynb notebook ready for Google Colab to fine-tune your SLM of choice, powered by Unsloth. - Future: AI-driven dataset analysis to automatically choose optimal training strategies and hyperparameters. Stack: Python, Unsloth, Google Colab, LLM Site: https://tuneslm.himanshuat.com ## MockServer *ongoing · DevTools · Nov 2025 -- ongoing* - A CLI tool that spins up a mock backend so you can build the frontend before the real one exists, with a GUI to configure it. - Inject latency and errors on demand to test loading and error states in the frontend. Stack: TypeScript, Next.js, Express Site: https://mockserver.himanshuat.com Source: https://github.com/Himasnhu-AT/mockserver ## FastSearch *paused · Software · April 2024 -- June 2024* - A privacy-focused search engine written in Rust, as an alternative to Google Search. - Ranks results with TF-IDF over an inverted index built for fast lookups. Stack: Rust, JavaScript Source: https://github.com/Himasnhu-AT/FastSearch ## FallenAngel-SE/CY *archived · AI tools · 2024 -- Present [Paused because of funds issue]* - A suite of AI tools for development and security, built as a coding assistant. - Runs on the LLaMA 3.1 70B model. - Includes interactive learning features. Stack: AI, Python, TypeScript, Cybersecurity demo1: https://www.linkedin.com/posts/himanshu-at_testing-out-new-product-revealed-soon-activity-7232783015360221185-t-Dr demo2: https://www.linkedin.com/posts/himanshu-at_llm-really-matters-a-lot-in-my-earlier-activity-7233145551096012800-d87p demo3: https://www.linkedin.com/posts/himanshu-at_edtech-machinelearning-pythonprogramming-activity-7233733311326474241-FbZK ## Dynamo-Prisma *completed · package · 2024* - Generates Prisma models from a JSON schema, so you don't hand-write them. Stack: TypeScript, Prisma, Node.js Source: https://github.com/techsavvyash/dynamo-prisma ## ScratchAI-Algo *in-progress · Learning · SynthAiLabs · April 2024 -- May 2024* - AI/ML algorithms implemented from scratch, each with a written explanation. - Covers KNN, logistic regression, K-means, decision trees, SVM, and more; 10+ so far. - Built for the SynthAiLabs education work, open to contributions. Stack: Python, NumPy, Pandas Source: https://github.com/Himasnhu-AT/ScratchML-Algo ## QM - Research *in-progress · Learning · April 2024 -- June 2024* - A research repo on using quantum computing against encryption like RSA, SHA-256, and SHA-128. - Built quantum circuits in Qiskit: Quantum Fourier Transform, GHZ state, and commutation-relation experiments. - Wrote OpenQASM for quantum registers and entanglement experiments. Stack: Python, Qiskit, OpenQASM Source: https://github.com/Himasnhu-AT/QM-res ## RustyExplorer *completed · Application · Oct 2023 -- Dec 2023* - A cross-platform file explorer built with Rust and TypeScript on Tauri. - Searches files noticeably faster than Windows Explorer, with cached directory listings to keep navigation snappy. Stack: Rust, TypeScript, Tauri, Tailwind CSS, Vite, React Source: https://github.com/Himasnhu-AT/rusty-explorer ## Linux Mini-Firewall *completed · Learning · July 2023* - A mini firewall application that allows users to add, delete, and print rules based on criteria like source IP, destination IP, and ports. Stack: C Source: https://github.com/Himasnhu-AT/Linux_mini-Firewall ## indieOS *paused · Learning · May 2024* - A demo operating system written from scratch for the x86 architecture. - Includes a bootloader, shell, ports, and custom memory management. Stack: Assembly, C, Bash, OS Source: https://github.com/Himasnhu-AT/indieOS ## Codeforces Scraper *completed · scrapper · June 2024* - A scraper to extract problems from Codeforces HTML pages and return them in JSON format. Stack: Python Source: https://github.com/Himasnhu-AT/codeforce_scrapper ## Log Ingestor *archived · package · Aug 2024 -- Sep 2024* - A high-performance log ingestor that stores logs in a database or prints them to the terminal in a clean format. Stack: Rust Source: https://github.com/Himasnhu-AT/LogInjestor-depreciated- ## ZenStore (In-memory DB) *paused · Application · Oct 2024* - A minimalistic in-memory database inspired by Redis, built entirely in Rust. Stack: Rust Source: https://github.com/Himasnhu-AT/ZenStore ## frm (Fast Delete) *completed · package · Dec 2024* - A high-performance, safe alternative to the `rm` command for fast file deletion, written in Rust. Stack: Rust Source: https://github.com/Himasnhu-AT/frm ## roadmap.ai *archived · Application · Oct 2024 -- Nov 2024* - An AI-powered tool that generates personalized learning roadmaps for various technical topics. Stack: Next.js, AI Site: https://roadmap-ai-xi.vercel.app ## biv *completed · package · Mar 2025* - An experimental, high-performance binary data format designed as a faster alternative to JSON. Stack: Rust Source: https://github.com/Himasnhu-AT/biv ## ruspy *paused · Learning · Nov 2024 -- Mar 2025* - A hobby programming language combining Rust's efficiency with Python's simplicity, built to learn compiler design. Stack: Rust, Compiler Design Source: https://github.com/Himasnhu-AT/ruspy ## PoW-Shield (DDoS Prevention) *completed · Application · Aug 2023* - Prevents DDoS attacks using a Proof of Work system to validate requests before forwarding them to the server. Stack: C++, Rust, C, TypeScript, Shell Source: https://github.com/Himasnhu-AT/PoW-Shield ## git-history *completed · package · Sep 2024* - A utility to convert a repository's entire Git commit history into a structured JSON file. Stack: Rust, Python Source: https://github.com/Himasnhu-AT/git-history ## localai *completed · Application · Apr 2025* - A desktop app to chat with and compare various local AI models deployed via Ollama, with an integrated web search tool. Stack: Tauri, Rust, Next.js, Ollama Site: https://github.com/Himasnhu-AT/localai ## draw (Excalidraw for Mac) *completed · Application · Feb 2025 -- Mar 2025* - A native macOS and tablet application that replicates the functionality of Excalidraw for offline use. Stack: Tauri, Rust, Next.js Source: https://github.com/Himasnhu-AT/draw ## Syntaxia (learncode) *paused · Application · Apr 2025* - An interactive desktop app for learning to code through hands-on practice and AI-driven guidance. Stack: Tauri, Rust, Next.js, Ollama, Gemini API Source: https://github.com/Himasnhu-AT/learncode ## dbquery *completed · Application · July 2025* - An AI agent that translates natural language queries into SQL commands to interact with databases. Stack: Tauri, Rust, Next.js, Ollama, Gemini API Source: https://github.com/Himasnhu-AT/dbquery ## codehaven.org *in-progress · Application · Nov 2024 -- Dec 2024* - A minimalistic, memory-friendly IDE built with Tauri, inspired by Neovim, as a GUI-based alternative to VSCode. Stack: Tauri, Rust, Next.js, Ollama Source: https://github.com/Himasnhu-AT/codehaven.org --- Source: https://himanshuat.com/projects ======================================================================== # Skills ## Languages - **TypeScript** — Daily driver for frontend, Next.js apps, CLI tools, and most production services. - **Python** — Go-to for ML, data pipelines, scraping, and quick experiments. - **Rust** — Reach for it when latency and memory matter — inference platforms, CLIs, file tools. - **Swift** — Native iOS/macOS work — SwiftUI for clean app surfaces. - **Dart** — Cross-platform mobile builds via Flutter. - **C/C++** — Systems-level work, OS experiments, performance-critical bits. - **Go** — Backend services and concurrent network tooling. - **Assembly x86** — Bootloaders and kernel work in indieOS. ## Frontend - **React** — Default UI library for every web product I ship. - **Next.js** — Full-stack framework of choice — App Router, RSC, edge-ready. - **TailwindCSS** — Styling layer for every Next.js / React surface I build. - **SwiftUI** — Declarative UI for native macOS and iOS apps. - **Flutter** — Cross-platform mobile when Swift isn't the right tool. - **Tauri** — Lightweight Rust-backed desktop shell — used for RustyExplorer, draw, dbquery, codehaven. ## Backend - **Nest.js** — Structured Node.js backend framework — used at scale for SynthAiLabs services. - **Flask** — Lightweight Python web framework for quick AI inference endpoints. - **http (rust)** — Low-level HTTP primitives in Rust for custom servers. - **Axum** — Async Rust web framework powering the AI Inference Platform. - **Tokio (rust)** — Async runtime underpinning every Rust service I write. - **Node.js** — Backend runtime for TypeScript services, CLIs, and dev tooling. - **FastAPI** — Async Python framework for ML serving and rapid API prototyping. - **Express** — Minimal Node.js framework — used in MockServer's mock-engine. - **Prisma** — Type-safe ORM for PostgreSQL-backed Next.js apps. ## Database - **PostgreSQL** — Primary relational store for production apps. - **MongoDB** — Document store for flexible-schema workloads. - **TimescaleDB** — Time-series Postgres extension for metrics and event data. - **Milvus** — Vector database for high-recall semantic search at scale. - **Redis** — Cache, message queue, and L2 layer for the AI Inference Platform. - **MySQL** — Used in legacy/contract setups where Postgres isn't the default. - **pgvector** — Embedding storage and similarity search inside Postgres. - **OpenSearch** — Powered low-latency profile retrieval for the WhatsApp AI agents at Tejas AI. ## DevOps & Cloud - **Docker** — Default packaging for every backend service I deploy. - **Kubernetes** — Container orchestration for horizontally scaled inference workloads. - **AWS** — Primary cloud — EC2, S3, RDS, Lambda for production deploys. - **Azure** — Used selectively for client engagements requiring Microsoft stack. - **GCP** — BigQuery, Cloud Run, and Vertex AI when the workload fits. - **ShaktiCloud** — India-hosted GPU cloud — used at AuraX for diffusion inference. - **CI/CD** — Build pipelines, automated testing, and zero-touch deploys. - **GitHub Actions** — Workflow engine of choice — covers every repo I ship. - **Prometheus** — Metrics scraping for the AI Inference Platform observability stack. - **Grafana** — Dashboards on top of Prometheus for latency, cache, and model health. ## AI/ML/QM - **PyTorch** — Core deep-learning framework for VTO research and custom models. - **TensorFlow** — Used selectively where the ecosystem fits (Keras, TFLite). - **scikit-learn** — Classical ML for baselines and feature engineering. - **diffusers** — HuggingFace library used heavily for Flux-based VTO pipelines. - **CUDA** — Custom kernels and inference tuning on NVIDIA GPUs. - **Unsloth** — Memory-efficient fine-tuning library — powers TuneSLM. - **Qiskit** — IBM's quantum SDK — used for entanglement and QFT experiments. - **OpenQASM** — Quantum assembly for circuit-level optimization in QM research. - **LLM** — General large language model stack — OpenAI, Anthropic, Gemini, local Ollama. - **LoRA** — Low-rank adapters — central to AuraX's C-AAC composition algorithm. - **Ollama** — Local LLM runtime — used in localai, codehaven, and dbquery. - **Gemini API** — Google's LLM API — used in dbquery and Syntaxia. - **OpenAPI** — Spec-driven backend tooling — powers OpenCrow's server-side tool layer. - **Mem0** — Persistent conversational memory layer built for the Tejas AI WhatsApp agents. --- Source: https://himanshuat.com/skills ======================================================================== # Changelog ## Started a systems measurement study on mmap versus read() *2026-04-01 · research* Re-measuring the CIDR'22 result that mmap underperforms read() for large scans, on 16 KiB-page Apple Silicon, which that literature does not cover. Still in progress: bare-metal Linux confirmation is outstanding. Tags: Systems, Performance, Rust, Research Link: [object Object] ## Released MockServer v1.3.0 *2025-12-14 · project* Released version 1.3.0 of MockServer. Tags: npm package, mock server, release Link: [object Object] ## Wound down AuraX *2025-12-01 · career* Closed AuraX after a year. The research held up, reaching FID 18.5 and SSIM 0.85 on complex global garments, and we had signed deals with Aditya Birla Fashion Group and Mensa Brands labels. Larger players entered the space and the runway ran short. Tags: AuraX, Startup ## Released MockServer *2025-11-22 · project* Generate realistic mock data, simulate network chaos, and build robust UIs with a single CLI command to run backend. Tags: npm package, mock server Link: [object Object] ## Released Fintracker *2025-11-15 · project* Published my first Flutter application to the Play Store. Fintracker is a privacy-first expense manager that keeps all data local. Built with Flutter and Hive. Tags: Flutter, Mobile, Release Link: [object Object] ## Co-founded AuraX, joined as CTO *2024-12-01 · career* Co-founded AuraX, a fashion-tech startup incubated at IIIT-Hyderabad's CIE, leading research on diffusion models for virtual try-on and building the GPU inference infrastructure behind it. Tags: AuraX, Diffusion Models, CTO --- Source: https://himanshuat.com/changelog ======================================================================== # Writing 48 posts, newest first. Append `.md` to any post URL for its Markdown source. - **The backtest is the experiment, and most of them aren't controlled** (August 12, 2026) — Self-taught quant research mostly fails on bias-laden backtests rather than on the math. So I front-loaded the pitfall literature before running anything, and it changed what I built first — infrastructure, not strategies. https://himanshuat.com/blogs/the-backtest-is-the-experiment.md - **554 million ticks and the wrong reason mmap loses** (August 12, 2026) — Everyone repeats that mmap is bad for large scans. On Apple Silicon that turned out to be true, false, and true again depending on working-set size — and the mechanism everybody names for it is not the mechanism doing the work. https://himanshuat.com/blogs/the-fault-policy-not-the-page-size.md - **LLM as judge, what the numbers actually say** (June 23, 2026) — I let a judge gate whether changes shipped, because the blog posts said it agreed with humans about 80% of the time. Six weeks in, I reported an improvement that was the judge flipping on rerun. This is what that 80% actually measures, and how little of it survives chance correction. https://himanshuat.com/blogs/llm-as-judge-what-the-numbers-actually-say.md - **Loop engineering** (June 02, 2026) — Agents rarely fail because the model can't do the task. They fail because the same task succeeds on Monday and fails on Thursday while somebody is watching, and nothing in the loop notices. Fixing that meant engineering the loop instead of the prompt. https://himanshuat.com/blogs/loop-engineering.md - **Graph agents and when the graph earns its keep** (May 12, 2026) — A graph framework hands you a second control flow language layered on top of Python, and you pay for it in state design, checkpoint size, and debugging. I built the same agent both ways under constraints I couldn't negotiate. The line where the graph wins turned out to be narrow, and most of my systems fell on the other side of it. https://himanshuat.com/blogs/graph-agents-when-the-graph-earns-its-keep.md - **When the documents disagree** (March 17, 2026) — An answer cited a specification section by number and was wrong, because an addendum had already revised that clause and nothing in the index knew it. Chasing that led to a FinanceBench result where giving each document its own vector store moved GPT-4-Turbo from 19% to 50% with no change to the model. https://himanshuat.com/blogs/when-the-documents-disagree.md - **Deterministic retrieval, and the parts of RAG that refuse to sit still** (February 17, 2026) — I asked a resume search for candidates with five years of experience and got someone with three. Embeddings encode semantic proximity, not predicate satisfaction. A worked account of hybrid SQL plus vector retrieval, the post-filter trap in pgvector, and why determinism is a property of the whole stack rather than a temperature setting. https://himanshuat.com/blogs/deterministic-rag.md - **Database agents that know when they're wrong** (January 20, 2026) — A SQL agent that errors is a nuisance. A SQL agent that returns a clean, confident, wrong number is a liability. Notes from building the scaffolding around text-to-SQL and document-store querying, and on why the benchmarks everyone quotes are partly measuring annotation noise. https://himanshuat.com/blogs/db-agents-that-know-when-theyre-wrong.md - **Production-Grade RAG: A Blueprint for Scalable, Real-Time Architecture** (November 27, 2025) — A RAG system fails quietly. It answers from a document that no longer exists, retrieval scores slide for weeks, and nothing in the trace is an error. This is the architecture I'd build to keep one fresh, fast, and observable. https://himanshuat.com/blogs/production-grade-rag-architecture-blueprint.md - **The Self-Correcting RAG: Implementing Agentic and Recursive Retrieval Loops** (November 18, 2025) — Static RAG retrieves once and hopes the first search was enough. This is the loop version, where the model judges its own context, names what is missing, writes a new query, and searches again until it stops or hits the cap. TypeScript, raw APIs, ideas from FLARE and Self-RAG. https://himanshuat.com/blogs/self-correcting-rag-agentic-recursive-retrieval.md - **Precision Tuning: Optimizing the Retriever vs. the Generator in Your RAG Pipeline** (November 06, 2025) — The wrong chunk comes back ranked first and the generator answers on top of it anyway. Fixing that means deciding which half of the pipeline is actually broken, the retriever or the generator, and then paying for the right one in the right order. https://himanshuat.com/blogs/optimizing-rag-fine-tuning-retriever-vs-generator.md - **Beyond Naive RAG: Advanced Chunking and Embedding Strategies for Superior Retrieval** (October 28, 2025) — A fixed-size chunker splits the sentence that answers the query across two chunks, and neither half retrieves. The fact is in your corpus and the system says it doesn't know. Chunking strategy and embedding choice are the two levers that fix that, and Hit Rate and MRR are how you prove it. https://himanshuat.com/blogs/advanced-rag-chunking-embedding-strategies.md - **Building Your First RAG System: From Zero to QA Hero** (October 14, 2025) — Ask a model about a document it never trained on and it will answer anyway, confidently and wrongly. Retrieval fixes that by handing the model the document at inference time instead of trying to train it in. This post builds the whole pipeline from scratch — loading, chunking, embedding, vector search, and the prompt that is allowed to say it doesn't know. https://himanshuat.com/blogs/building-first-rag-system-from-zero-to-qa-hero.md - **We beat the benchmarks and ran out of runway** (September 25, 2025) — Flux-VTON+ scores 0.85 SSIM on global traditional garments where SDXL manages 0.45, and we have signed deals with real apparel groups. We're winding the company down anyway. This is what I think the research was worth, written while both of those things are true at once. https://himanshuat.com/blogs/we-beat-the-benchmarks-and-ran-out-of-runway.md - **Three clouds, one inference path** (September 17, 2025) — Inference ran on three providers because one brand's images could not leave the country, one workload wanted a machine you could leave dirty, and the other sat at zero for four days and then arrived as a catalogue drop. This is the routing rule that made it survivable, and the bill I never measured. https://himanshuat.com/blogs/three-clouds-one-inference-path.md - **From a ComfyUI graph to an API a brand's team can call** (September 09, 2025) — Our try-on pipeline lived as a ComfyUI graph with implicit state, hand-picked seeds, and custom nodes tracking whatever was on main that week. Turning it into something another company's engineers could call meant writing down which code ran, which seed, when the work happens, and what we refuse. https://himanshuat.com/blogs/from-comfyui-graph-to-an-api.md - **Fitting a saree into 24GB: the context window trick** (September 02, 2025) — A saree needs pixel density to hold its pleats, and a full-frame high-resolution pass doesn't fit on a 24GB card. The fix was to stop treating the frame as the unit of inference and hand the model only the region that matters. https://himanshuat.com/blogs/fitting-a-saree-into-24gb.md - **FID was the wrong metric for the problem we had** (August 26, 2025) — Aggregate FID went 22.1 to 18.5 and I nearly read that as months of adapter work buying four points. Half the eval set was already close to solved and it dragged the average toward nothing. Splitting by garment category was the change that made the result visible. https://himanshuat.com/blogs/fid-was-the-wrong-metric.md - **Composing adapters that disagree with each other** (August 05, 2025) — Two LoRA experts merged with two scalars shipped fine. A whole library of them for demographics, poses, lighting and backgrounds did not, because adapters conflict per layer and one blend weight per adapter cannot express that. https://himanshuat.com/blogs/conflict-aware-adapter-composition.md - **What happens when you merge two LoRAs** (July 15, 2025) — Two adapters that each pass on their own can make each other worse the moment you sum them into the base weights. We shipped a hand-tuned coefficient pair because we had to ship something, and months later I still don't have a principled answer for how those coefficients should be chosen. https://himanshuat.com/blogs/what-happens-when-you-merge-two-loras.md - **Serverless GPUs for bursty training** (June 17, 2025) — Our finetuning load arrived in bursts, a few hard days then a quiet week, and a rented GPU box covers that badly. The split we ended up running, and the one number I still can't give you. https://himanshuat.com/blogs/serverless-gpus-for-bursty-training.md - **One LoRA per failure mode** (May 20, 2025) — Put a hand on a hip and our try-on pipeline painted the garment straight over the fingers. The obvious fix, one bigger dataset covering every hard case at once, quietly made the draping worse. Splitting the finetune into two adapters, one per named failure mode, is what actually worked. https://himanshuat.com/blogs/one-lora-per-failure-mode.md - **Generic aesthetic scorers hate e-commerce photography** (April 29, 2025) — We pointed an open aesthetic scorer at a batch of generated product shots and it ranked the moody, low-key ones highest. That's a defensible opinion about photographs and the wrong answer for a catalogue. This is the scorer we built instead, and the parts of it I can't put a number on. https://himanshuat.com/blogs/teaching-a-model-what-a-catalogue-looks-like.md - **LoRA finetuning: the hyperparameters settled in a week, the data never did** (April 08, 2025) — An early draping adapter kept painting the same flat plane across the back of the shoulder, on prompts with nothing in common. That is memorisation, not physics, and it traced back to one shoot. The config for these LoRAs settled after a handful of runs and never moved again. Everything that improved after that came out of the dataset. https://himanshuat.com/blogs/lora-finetuning-what-moved-the-needle.md - **Two streams: why one diffusion pass couldn't do virtual try-on** (March 18, 2025) — Masking a torso and inpainting a garment demos well in an afternoon. Getting back the exact garment, with its weave and its print placement intact, is a different problem. One conditioning stream could not hold both the body and the reference, so Flux-VTON+ runs two. https://himanshuat.com/blogs/flux-fill-and-redux-two-streams.md - **What 5,000 images taught me about curation** (February 25, 2025) — More data made the model worse in a way no automatic metric could see. The seed adapters behind AuraX-V1 ended up trained on 5,000 images a person had looked at one by one. Here is what the filter caught, and what it kept missing. https://himanshuat.com/blogs/what-5000-images-taught-me-about-curation.md - **ComfyUI as a research lab** (February 04, 2025) — A custom node pack updated, a mask node changed its default on inverted masks, and our outputs moved with nothing changed on our side. Running a try-on diffusion pipeline as a node graph made every experiment a commit, and made very clear where a graph stops being a program. https://himanshuat.com/blogs/comfyui-as-a-research-lab.md - **Virtual try-on is not image generation** (January 14, 2025) — A text-to-image model invents a plausible shirt. Try-on has to reproduce the exact shirt the customer is already looking at, on a person whose face has to survive the process intact. Getting that wrong on a saree is the failure that started all of this. https://himanshuat.com/blogs/virtual-try-on-is-not-image-generation.md - **Part 5: The Architectural Frontier - Mamba, RAG, and the Future Beyond Attention** (June 07, 2024) — Attention costs O(N^2), so the context window is capped by arithmetic rather than by anything the model can or cannot do. Mamba trades the full history for a fixed-size state, and RAG moves the knowledge out of the weights. Here is what each one buys and what it costs. https://himanshuat.com/blogs/understanding-transformers-part-5-future-architectures.md - **Part 4: An Architectural Deep Dive - Why BERT and GPT Are Different Beasts** (May 31, 2024) — BERT and GPT run on the same self-attention machinery. One triangular matrix of negative infinities decides whether a model reads or writes, and every other difference between them follows from it. https://himanshuat.com/blogs/understanding-transformers-part-4-architecture-bert-vs-gpt.md - **Part 3: The Scaling Problem - Optimizing Transformer Memory and Compute** (May 24, 2024) — At N=8192 the attention score matrix is 67 million elements, about 268 MB per head per layer in float32. The math is fine. The memory traffic is what kills you, and FlashAttention, quantization, and LoRA each attack a different part of that bill. https://himanshuat.com/blogs/understanding-transformers-part-3-optimization-and-scaling.md - **Part 2: Assembling the Full Architecture - From Attention to a Working Model** (May 17, 2024) — Attention on its own is a lookup. A model is what you get when you wire lookups into a residual stream — positional encoding, masking, cross-attention — and check the tensor shape at every hop. https://himanshuat.com/blogs/understanding-transformers-part-2-full-architecture.md - **Part 1: The Attention Mechanism - Building the Core of the Transformer** (May 10, 2024) — An RNN has to compress the whole input sequence into one fixed-size vector before it is allowed to predict anything. Attention drops that constraint. This post builds it from raw tensor operations — projections, scaled dot product, softmax, weighted sum — in NumPy first, then PyTorch. https://himanshuat.com/blogs/understanding-transformers-part-1-attention-from-scratch.md - **Part 3: Scaling LangGraph - State Persistence, Checkpointing, and Parallelism** (May 03, 2024) — A long-running graph crashes and loses the reasoning chain that was mid-flight, because nothing in it was ever written down. Checkpointers and Pregel supersteps fix two different halves of that, and neither one gives you the parallelism you probably think you are buying. https://himanshuat.com/blogs/scaling-langgraph-persistence-and-parallelism.md - **Part 5: Production-Ready Agents - Implementing Human-in-the-Loop Supervision** (April 26, 2024) — An agent that can execute a $100,000 transfer needs a state that means wait. This is a small TypeScript orchestrator where the pause is a persisted state rather than an exception, and resuming means writing the human's decision into state before the loop runs again. https://himanshuat.com/blogs/production-ready-langgraph-human-in-the-loop.md - **Production-Grade Agent Architecture: Implementing Long-Term Memory and Human-in-the-Loop** (April 19, 2024) — A stateless agent forgets every decision the moment the process dies, and an agent holding a deploy tool will ship to production without asking anyone. Two pieces of engineering close both gaps: semantic memory over a vector store, and an approval gate that lives in the graph's state instead of in a callback. https://himanshuat.com/blogs/production-grade-agent-architecture-memory-human-in-the-loop.md - **Optimizing Agent Reliability: Debugging Trajectories and Prompt Engineering** (April 12, 2024) — An agent fails on the shape of its own output long before it fails on reasoning. Reading trajectories and writing the prompt as a format contract is what turns something that mostly works into something that holds up run after run. https://himanshuat.com/blogs/optimizing-langchain-agent-reliability-debugging-prompt-engineering.md - **LangChain Agents 101: Building Your First Autonomous Tool-User From Scratch** (April 05, 2024) — An agent is a prompt, a regex, a dict lookup, and a for loop. This builds one out of `requests` and `re` so that all four are yours to read, and names the places each of them breaks. https://himanshuat.com/blogs/langchain-agents-101-building-first-autonomous-tool-user.md - **Part 1: Beyond Sequential Chains - Your First Agentic Workflow with LangGraph** (March 29, 2024) — An LCEL chain that asks a clarifying question has nowhere to put the answer. It already ran to completion. LangGraph fixes that by making state explicit and letting the graph loop back on itself. https://himanshuat.com/blogs/getting-started-with-langgraph-agentic-workflows.md - **Beyond Pre-builts: Crafting Custom Tools for Domain-Specific LangChain Agents** (March 22, 2024) — A generic pre-built tool makes the agent guess the function, the arguments, and their shape, and guessing is where hallucinations come from. Custom tools replace the guess with a typed contract the model can read. https://himanshuat.com/blogs/crafting-custom-tools-for-langchain-agents.md - **Part 2: Building a Self-Correcting RAG with Conditional Edges** (March 15, 2024) — A retrieve-then-generate chain has no step whose job is to disagree with the previous one. This is what it takes to add that step — graph state, a critic node, and conditional edges that route the run back to retrieval when the answer is thin. https://himanshuat.com/blogs/building-self-correcting-rag-with-langgraph.md - **Part 4: Architecting Agent Teams - Hierarchical Workflows and Graph Composition** (March 08, 2024) — The manager graph needs a conditional edge, and the engine's edge type has no branch for a function. That gap is where a multi-agent system actually gets built - an explicit state contract, a small graph engine, and one node that invokes another graph. https://himanshuat.com/blogs/architecting-multi-agent-teams-with-langgraph.md - **From Solo Agent to Team Player: Architecting Multi-Agent Systems with LangGraph** (March 01, 2024) — One model asked to research, draft, review, and revise in a single pass does all four badly. Splitting the job across specialized agents and wiring them into a state machine graph makes every handoff explicit, and turns the supervisor into the only prompt that really matters. https://himanshuat.com/blogs/architecting-multi-agent-systems-with-langgraph.md - **StyleX- Replacement of TailWindCSS?** (December 16, 2023) — A look at StyleX, Facebook's CSS-in-JS library, and how it compares to Tailwind. https://himanshuat.com/blogs/stylex-new-frontend-styling-framework.md - **Setting up Linters, Syntax Highlighter, NerdTree, and More** (September 12, 2023) — A guide on configuring Neovim as a full-blown IDE. https://himanshuat.com/blogs/my-nvim-setup.md - **10 Essential Programming Concepts Every Developer Should Know** (February 11, 2023) — The core programming concepts that come up in almost every codebase. https://himanshuat.com/blogs/essential-programming-concepts.md - **The Promise and Potential of Quantum Computing** (February 05, 2023) — An overview of quantum computing, its principles, applications, and future prospects. https://himanshuat.com/blogs/quantum-computing-simplified.md - **A Beginner's Guide to Linux From Scratch** (February 16, 2021) — Learn how to build a Linux system from scratch without ISO images. https://himanshuat.com/blogs/A-Beginnner-Guide-to-Linux-From-Scratch.md --- Source: https://himanshuat.com/blogs ======================================================================== # Books 35 books with notes. - **80/20 Your Life!** — Damon Zahariades (2018): Damon Zahariades takes the Pareto principle out of the boardroom and aims it at ordinary life — arguing that a ruthless focus on the vital few, and the nerve to drop the trivial many, is what actually buys back your time. https://himanshuat.com/books/80-20-your-life.md - **Eat, Pray, Love** — Elizabeth Gilbert (2006): A memoir-as-self-help disguised as a travelogue: one year, three countries, one woman trying to put a divorced self back together without lying to herself about how it's going. https://himanshuat.com/books/eat-pray-love.md - **A Clockwork Orange** — Anthony Burgess (1962): A teenage thug, a state-mandated cure, and the most uncomfortable question in moral psychology: is a person who can no longer choose evil still a moral being at all? https://himanshuat.com/books/a-clockwork-orange.md - **The Stranger** — Albert Camus (1942): A man buries his mother, kills an Arab on a beach, and faces a court that wants him to lie about both. The novel that gave existentialism its case file. https://himanshuat.com/books/the-stranger.md - **The Body Keeps the Score** — Bessel van der Kolk (2014): A psychiatrist's three-decade investigation into how trauma reshapes the body and brain, and the surprisingly effective treatments that follow once we take that seriously. https://himanshuat.com/books/the-body-keeps-the-score.md - **The Death of Ivan Ilyich** — Leo Tolstoy (1886): A 19th-century judge realizes, over the course of a hundred quiet pages, that his successful life has been a costume — and that no one around him is willing to admit it with him. https://himanshuat.com/books/the-death-of-ivan-ilych.md - **Essays in Love** — Alain de Botton (1993): A philosopher's debut — half novel, half essay — that turns a single mostly-unremarkable relationship into a careful x-ray of how love actually works in the mind. https://himanshuat.com/books/essays-in-love.md - **Moonwalking with Einstein** — Joshua Foer (2011): A journalist follows the U.S. Memory Championship for a year, gets sucked in, and ends up winning it — using techniques that have been around for two thousand years. https://himanshuat.com/books/moonwalking-with-einstein.md - **The Celestine Prophecy** — James Redfield (1993): A divisive 1990s bestseller that smuggled a self-help manual inside a Peruvian adventure novel — and got a lot of people to take "coincidences" seriously for the first time. https://himanshuat.com/books/the-celestine-prophecy.md - **On the Shortness of Life** — Seneca (49): A two-thousand-year-old letter from a Stoic to a friend, arguing that life is not actually short — we just spend most of it not paying attention. https://himanshuat.com/books/on-the-shortness-of-life.md - **The Sense of an Ending** — Julian Barnes (2011): A man in his sixties is forced to revisit a single decision from his youth, and discovers his memory has been editing the record for forty years. https://himanshuat.com/books/the-sense-of-an-ending.md - **Meditations** — Marcus Aurelius (180): The private notebook of a Roman emperor talking himself, every morning for years, into being a slightly better human being. https://himanshuat.com/books/meditations.md - **Flowers for Algernon** — Daniel Keyes (1966): A first-person diary of a man whose IQ is surgically tripled — and then, slowly, slips back. The most affecting case study ever written about what intelligence actually costs. https://himanshuat.com/books/flowers-for-algernon.md - **The Analects** — Confucius (-500): Two and a half millennia of saved fragments from a teacher who refused to write a book, organized by his students into the most quietly demanding ethical handbook in print. https://himanshuat.com/books/the-analects.md - **The Republic** — Plato (-375): Plato's Socratic dialogue on justice, the ideal city, and the philosopher-king — the seed text most Western political philosophy is still arguing with. https://himanshuat.com/books/the-republic.md - **The Trial** — Franz Kafka (1925): A bank clerk is arrested by an opaque authority for a crime no one will name. Kafka's most-quoted study of bureaucratic paranoia and the dissolution of the self. https://himanshuat.com/books/the-trial.md - **The Brothers Karamazov** — Fyodor Dostoevsky (1880): Dostoevsky's last novel — four brothers, a murdered father, and the longest sustained meditation in fiction on faith, doubt, and what people do with their guilt. https://himanshuat.com/books/the-brothers-karamazov.md - **The Magic Mountain** — Thomas Mann (1924): Seven years inside a Swiss tuberculosis sanatorium where time stops and Europe argues with itself — Mann's portrait of the world right before it broke. https://himanshuat.com/books/the-magic-mountain.md - **Anthem** — Ayn Rand (1938): A short dystopia about a society that has erased the word "I" — and the man who, against the rules of his civilization, has to rediscover what it means. https://himanshuat.com/books/anthem.md - **We** — Yevgeny Zamyatin (1924): The Soviet dystopia that Orwell and Huxley both read before they wrote theirs — a glass-walled state where every minute is scheduled and individuality is a disease. https://himanshuat.com/books/we.md - **The Handmaid's Tale** — Margaret Atwood (1985): Atwood's near-future theocracy where women's bodies have been re-classified as state property — built, she's said, only from things that have actually happened somewhere. https://himanshuat.com/books/the-handmaids-tale.md - **Invisible Man** — Ralph Ellison (1952): An unnamed Black narrator describes the long process of becoming invisible to a country that refuses to see him — and what living underground does to the mind. https://himanshuat.com/books/invisible-man.md - **The Age of Reason** — Jean-Paul Sartre (1945): Two days in the life of a Parisian philosophy teacher trying to scrape together the money for an abortion — Sartre's existentialism with the abstractions cut out. https://himanshuat.com/books/the-age-of-reason.md - **Animal Farm** — George Orwell (1945): A barnyard fable about the Russian Revolution that ends up being a more general theory of how revolutions eat their own slogans. https://himanshuat.com/books/animal-farm.md - **The Screwtape Letters** — C.S. Lewis (1942): A senior demon writes to his nephew with advice on how to corrupt a single human soul — an inverted moral handbook that's funnier and more uncomfortable than it sounds. https://himanshuat.com/books/the-screwtape-letters.md - **The Strange Case of Dr. Jekyll and Mr. Hyde** — Robert Louis Stevenson (1886): The original split-self novella — a respectable Victorian doctor distills the part of himself he can't admit to, and lets it out for the evening. https://himanshuat.com/books/jekyll-and-hyde.md - **No Exit and Three Other Plays** — Jean-Paul Sartre (1944): Four short plays including the one with the famous line — "hell is other people" — though the line means something stranger in context than the way it's usually quoted. https://himanshuat.com/books/no-exit.md - **Man's Search for Meaning** — Viktor Frankl (1946): A psychiatrist's account of surviving Auschwitz, followed by a sketch of logotherapy — the school of psychology he built around the observation that meaning is something a person can keep when everything else is taken. https://himanshuat.com/books/mans-search-for-meaning.md - **Hunger** — Knut Hamsun (1890): An unnamed narrator wanders nineteenth-century Christiania, slowly starving — one of the first novels to take the inside of a single hungry mind as its entire subject. https://himanshuat.com/books/hunger.md - **Prometheus Rising** — Robert Anton Wilson (1983): A cult-favourite synthesis of Leary's eight-circuit model, Korzybski's general semantics, and Wilson's own scepticism — a guidebook to noticing your own mental shortcuts. https://himanshuat.com/books/prometheus-rising.md - **Wind, Sand and Stars** — Antoine de Saint-Exupéry (1939): The author of *The Little Prince* writing as himself — a pilot's memoir of mail routes over the Sahara, near-fatal crashes, and what those silences taught him about being human. https://himanshuat.com/books/wind-sand-and-stars.md - **Life is Elsewhere** — Milan Kundera (1969): Kundera's portrait of a young poet who treats his own life as a draft of someone else's poem — and the slow damage that does to the people around him. https://himanshuat.com/books/life-is-elsewhere.md - **Zorba the Greek** — Nikos Kazantzakis (1946): A bookish narrator meets an old miner named Zorba and spends a year being patiently dismantled by a man who refuses to live in his head. https://himanshuat.com/books/zorba-the-greek.md - **Factfulness** — Hans Rosling (2018): Ten instincts that quietly distort how educated people read the news — written by a Swedish physician who spent decades teaching the world to think in distributions, not headlines. https://himanshuat.com/books/factfulness.md - **How Doctors Think** — Jerome Groopman (2007): A Harvard internist on the cognitive shortcuts that get patients misdiagnosed — anchoring, availability, attribution — and how to ask a doctor the question that breaks the loop. https://himanshuat.com/books/how-doctors-think.md --- Source: https://himanshuat.com/books ======================================================================== # Work with me Available now — left Tejas AI (YC W25) in August 2026. The fastest way to work out whether I am useful to you is a conversation. Fifteen minutes, no deck. Book a slot: https://cal.com/himanshuat/makecustom12?duration=15 ## What I am open to - **Full-time** — AI infrastructure, search and retrieval, or quantitative research. - **Contract** — scoped work on retrieval, agent reliability, or inference cost. - **Research** — collaborations; currently on measurement studies in systems and finance. ## What I can do ### Search and retrieval Hybrid keyword and vector search that returns the thing someone actually asked for. Structure-aware chunking, BM25 fused with dense retrieval, metadata filters, and a latency-budgeted rerank on top — built and judged against a real query set rather than a vibe. _Built on OpenSearch at Tejas AI over a corpus indexed from ~60 ATS platforms._ ### Agents that survive production A router with specialist sub-agents, routing decided in code rather than left to the model's discretion, and memory that persists across sessions instead of living in a context window. Plus the unglamorous half: failure taxonomies from real traces and eval sets wired into CI. _The WhatsApp agent behind Elip, with state across Redis, Postgres and S3._ ### Inference infrastructure and cost Complexity-based routing so easy queries stop hitting a frontier model, two-tier semantic caching so paraphrased repeats never reach the API, and enough observability to prove the bill actually moved. _An open-source inference platform in Rust — L1 Moka, L2 Redis, Prometheus and Grafana._ ### Applied research, taken to production Diffusion pipelines, LoRA adapter composition, and the measurement work that tells you whether a result is real. I care about the part after the paper: making it run, making it cheap enough, and making the numbers hold up when someone else checks. _Flux-VTON+ at AuraX — research through to production inference and signed commercial deals._ ## Contact - Book 15 minutes: https://cal.com/himanshuat/makecustom12?duration=15 - Email: hyattherate2005@gmail.com --- Source: https://himanshuat.com/hire ======================================================================== # Knowledge graph An interactive force-directed graph of everything on this site — 111 nodes and 250 edges connecting projects, research, roles and skills. The rendered page is a canvas; this is the same structure as text. ## Root (1) - **~/himanshu** (root): Knowledge graph of everything I've built — projects, research, companies, and the skills threading them together. Click any node to inspect it. ## Projects (33) - **#projects** (domain): Personal and OSS projects, grouped by category. - **#ai** (hub): Project category: ai. - **CCProxy** (project): (Mar 2026 -- ongoing) A universal proxy that lets Claude Code talk to any AI provider: OpenAI, Gemini, or local Ollama models. — https://ccproxy.himanshuat.com/ - **OpenCrow** (project): (Jan 2026 -- ongoing) Open-source AI agent platform for embedding intelligent customer support assistants into web applications. Features an embeddable React chat widget with dual-layer tooling: server-side tools via OpenAPI specs and client-side tools executed in the browser. Dynamic tool merging lets the LLM pick between server and client actions per request. Includes a management dashboard for agents, products, and integrations. Inspired by usecrow.ai. — https://opencrow.himanshuat.com - **#infra** (hub): Project category: infra. - **AI Inference Platform** (project): (Dec 2025 -- Jan 2026) High-performance AI inference platform with intelligent model routing — classifies query complexity via embeddings and heuristics, routing simple queries to fast/cheap models and complex ones to GPT-4 class models. Multi-level caching: L1 in-memory (Moka) + L2 distributed (Redis) with semantic similarity matching via cosine similarity on embeddings. Full observability stack with Prometheus metrics and pre-configured Grafana dashboards for inference latency, cache hit rates, and model health. Fully async Rust architecture with horizontal scaling via Docker and Kubernetes. — https://infer.himanshuat.com - **TuneSLM** (project): (Jan 2026 -- ongoing) Drop your dataset and answer a few questions — get an ipynb notebook ready for Google Colab to fine-tune your SLM of choice, powered by Unsloth. Future: AI-driven dataset analysis to automatically choose optimal training strategies and hyperparameters. — https://tuneslm.himanshuat.com - **#devtools** (hub): Project category: devtools. - **MockServer** (project): (Nov 2025 -- ongoing) A CLI tool that spins up a mock backend so you can build the frontend before the real one exists, with a GUI to configure it. Inject latency and errors on demand to test loading and error states in the frontend. — https://mockserver.himanshuat.com - **FastSearch** (project): (April 2024 -- June 2024) A privacy-focused search engine written in Rust, as an alternative to Google Search. Ranks results with TF-IDF over an inverted index built for fast lookups. — https://github.com/Himasnhu-AT/FastSearch - **FallenAngel-SE/CY** (project): (2024 -- Present [Paused because of funds issue]) A suite of AI tools for development and security, built as a coding assistant. Runs on the LLaMA 3.1 70B model. Includes interactive learning features. — https://www.linkedin.com/posts/himanshu-at_testing-out-new-product-revealed-soon-activity-7232783015360221185-t-Dr - **Dynamo-Prisma** (project): (2024) Generates Prisma models from a JSON schema, so you don't hand-write them. — https://github.com/techsavvyash/dynamo-prisma - **#learning** (hub): Project category: learning. - **ScratchAI-Algo** (project): (April 2024 -- May 2024) AI/ML algorithms implemented from scratch, each with a written explanation. Covers KNN, logistic regression, K-means, decision trees, SVM, and more; 10+ so far. Built for the SynthAiLabs education work, open to contributions. — https://github.com/Himasnhu-AT/ScratchML-Algo - **QM - Research** (project): (April 2024 -- June 2024) A research repo on using quantum computing against encryption like RSA, SHA-256, and SHA-128. Built quantum circuits in Qiskit: Quantum Fourier Transform, GHZ state, and commutation-relation experiments. Wrote OpenQASM for quantum registers and entanglement experiments. — https://github.com/Himasnhu-AT/QM-res - **#apps** (hub): Project category: apps. - **RustyExplorer** (project): (Oct 2023 -- Dec 2023) A cross-platform file explorer built with Rust and TypeScript on Tauri. Searches files noticeably faster than Windows Explorer, with cached directory listings to keep navigation snappy. — https://github.com/Himasnhu-AT/rusty-explorer - **Linux Mini-Firewall** (project): (July 2023) A mini firewall application that allows users to add, delete, and print rules based on criteria like source IP, destination IP, and ports. — https://github.com/Himasnhu-AT/Linux_mini-Firewall - **indieOS** (project): (May 2024) A demo operating system written from scratch for the x86 architecture. Includes a bootloader, shell, ports, and custom memory management. — https://github.com/Himasnhu-AT/indieOS - **Codeforces Scraper** (project): (June 2024) A scraper to extract problems from Codeforces HTML pages and return them in JSON format. — https://github.com/Himasnhu-AT/codeforce_scrapper - **Log Ingestor** (project): (Aug 2024 -- Sep 2024) A high-performance log ingestor that stores logs in a database or prints them to the terminal in a clean format. — https://github.com/Himasnhu-AT/LogInjestor-depreciated- - **ZenStore (In-memory DB)** (project): (Oct 2024) A minimalistic in-memory database inspired by Redis, built entirely in Rust. — https://github.com/Himasnhu-AT/ZenStore - **frm (Fast Delete)** (project): (Dec 2024) A high-performance, safe alternative to the `rm` command for fast file deletion, written in Rust. — https://github.com/Himasnhu-AT/frm - **roadmap.ai** (project): (Oct 2024 -- Nov 2024) An AI-powered tool that generates personalized learning roadmaps for various technical topics. — https://roadmap-ai-xi.vercel.app - **biv** (project): (Mar 2025) An experimental, high-performance binary data format designed as a faster alternative to JSON. — https://github.com/Himasnhu-AT/biv - **ruspy** (project): (Nov 2024 -- Mar 2025) A hobby programming language combining Rust's efficiency with Python's simplicity, built to learn compiler design. — https://github.com/Himasnhu-AT/ruspy - **PoW-Shield (DDoS Prevention)** (project): (Aug 2023) Prevents DDoS attacks using a Proof of Work system to validate requests before forwarding them to the server. — https://github.com/Himasnhu-AT/PoW-Shield - **git-history** (project): (Sep 2024) A utility to convert a repository's entire Git commit history into a structured JSON file. — https://github.com/Himasnhu-AT/git-history - **localai** (project): (Apr 2025) A desktop app to chat with and compare various local AI models deployed via Ollama, with an integrated web search tool. — https://github.com/Himasnhu-AT/localai - **draw (Excalidraw for Mac)** (project): (Feb 2025 -- Mar 2025) A native macOS and tablet application that replicates the functionality of Excalidraw for offline use. — https://github.com/Himasnhu-AT/draw - **Syntaxia (learncode)** (project): (Apr 2025) An interactive desktop app for learning to code through hands-on practice and AI-driven guidance. — https://github.com/Himasnhu-AT/learncode - **dbquery** (project): (July 2025) An AI agent that translates natural language queries into SQL commands to interact with databases. — https://github.com/Himasnhu-AT/dbquery - **codehaven.org** (project): (Nov 2024 -- Dec 2024) A minimalistic, memory-friendly IDE built with Tauri, inspired by Neovim, as a GUI-based alternative to VSCode. — https://github.com/Himasnhu-AT/codehaven.org ## Research (8) - **#research** (domain): Research output — VTO, MoE, quantum, transformers. - **It's the Fault Policy, Not the Page Size: Why mmap Loses to read() for Out-of-Core Scans on Apple Silicon** (research): Re-measuring mmap vs read() on Apple Silicon: a working-set crossover, attributed to OS fault-population policy rather than page size. — https://himanshuat.com/research/mmap-vs-read-apple-silicon - **Conflict-Aware Adapter Composition (C-AAC)** (research): Outperforms base Flux, Imagen 4, and ChatGPT in skin quality and aesthetic appeal. — https://himanshuat.com/research/c-aac - **Flux-VTON+: Hybrid Flux Inpainting and Multi-LoRA Expert Fusion** (research): State-of-the-art VTON pipeline synergizing Flux Fill and Redux with Multi-LoRA Fusion, achieving an 85% success rate on complex garments. — https://himanshuat.com/research/renu-virtual-try-on - **MoE: Brain-Inspired LLM Architecture** (research): A whitepaper advancing MoE as a computational analogue to the brain's functional specialization. — https://himanshuat.com/research/moe-brain - **High-Dimensional Lattice-Based Quantum Encryption (LWE)** (research): A PQC scheme using the LWE problem with detailed mathematical formulations. — https://himanshuat.com/research/lattice-based-quantum-encryption - **GPT-Neo: Transformer Implementation** (research): Implementing 'Attention is All You Need' from scratch. — https://himanshuat.com/research/gpt-neo - **Custom Neural Network for MNIST** (research): A custom-built Neural Net for the MNIST digit recognizer dataset. — https://himanshuat.com/research/custom-n-n-mnist ## Experience (7) - **#experience** (domain): Companies and roles across the timeline. - **Tejas AI** (company): Software Development Engineer (SDE) · Mar 2026 -- Aug 2026. Built Elip, an AI career coach on WhatsApp (with web, a Chrome extension, and a mobile app) that helps job seekers fix their resume, build a search profile, and get matched to jobs. The core was a FastAPI WhatsApp agent, a router model with specialist sub-agents, that parsed resumes and LinkedIn profiles, ran resume reviews, and returned ranked jobs, with cross-session memory across Redis, Postgres, and S3. Built the search layer over OpenSearch (hybrid keyword and vector search with Bedrock embeddings) and a fleet of Go services that index jobs from around 60 ATS and job boards. Built the Go backend behind the web and mobile apps, plus a Chrome extension that autofills multi-step job applications and tailors a resume to each posting. Stack: TypeScript, Go, Node.js, Nest.js, PostgreSQL, Redis, OpenSearch, Mem0, LLM, Docker, AWS. — https://himanshuat.com/experience - **AuraX** (company): AI Researcher & CTO · Dec 2024 -- Dec 2025. Co-founded AuraX, a fashion-tech startup funded by IIIT-Hyderabad's CIE incubator. Led research on diffusion models for virtual try-on (VTO), achieving quality that outperformed Google, Stable Diffusion, and leading Indian startups in benchmark comparisons. Signed commercial deals with Aditya Birla Fashion Group and multiple brands under Mensa Brands. Built the backend and deployment infrastructure on AWS and ShaktiCloud, and tuned inference latency for production. Wound the company down as bigger players entered the space and our runway ran short. Stack: Python, PyTorch, diffusers, LoRA, CUDA, LLM, FastAPI, PostgreSQL, Redis, Docker, AWS, ShaktiCloud, Next.js, Nest.js. — https://himanshuat.com/experience - **SynthAiLabs** (company): Founder · Dec 2023 -- Feb 2025. Researched Mixture-of-Experts (MoE) LLM architectures. Built learncode (interactive AI lessons) and eduAI (an agent that answers questions about course material). Grew techos.synthailabs to 11,000 registered students across India. Co-founded the OpenCodeDevelopers (OCD) society, which grew by word of mouth and helped 800+ students cut their CP/DSA ramp from six months to two. Stack: TypeScript, Next.js, Node.js, Express, MongoDB, PostgreSQL, Redis, Docker, CI/CD, React, TailwindCSS. — https://himanshuat.com/experience - **Tublian** (company): AI Dev · July 2024 -- Dec 2024. Set up Docker-based build and deployment workflows. Added code obfuscation to protect the source before public deployment. Contributed to TublianOS, working on performance and the user-facing experience. Completed a system-design masterclass that used Google's search architecture as its worked example. Open Source Collaborator | Project Mentor · Dec. 2023 -- April 2024. Contributed to several open-source projects and mentored 20+ newer developers. Helped build DevDocGenie, a RAG-based chatbot for searching documentation. Reverse-engineered features like Samsung's live call translation as build exercises. Stack: Python, LLM, Ollama, FastAPI, Docker, TypeScript, React, Node.js, GitHub Actions. — https://himanshuat.com/experience - **Stealth** (company): QM Researcher · March 2024 -- June 2024. [NDA], tested ML model deployment on quantum computers using qiskit Stack: Python, Qiskit, OpenQASM, C/C++. — https://himanshuat.com/experience - **Excelerate** (company): Project Manager · April 2023 -- Sept 2023. Led a team of six on a global education event with a $30,000 budget. Owned the project management docs, including the RACI matrix and risk register. Delivered on schedule. Stack: TypeScript, React, Next.js. — https://himanshuat.com/experience ## Skills (62) - **#skills** (domain): Languages, frameworks, infra, and AI/ML tools. - **#languages** (hub): 8 skills in Languages. - **TypeScript** (skill): Daily driver for frontend, Next.js apps, CLI tools, and most production services. - **Python** (skill): Go-to for ML, data pipelines, scraping, and quick experiments. - **Rust** (skill): Reach for it when latency and memory matter — inference platforms, CLIs, file tools. - **Swift** (skill): Native iOS/macOS work — SwiftUI for clean app surfaces. - **Dart** (skill): Cross-platform mobile builds via Flutter. - **C/C++** (skill): Systems-level work, OS experiments, performance-critical bits. - **Go** (skill): Backend services and concurrent network tooling. - **Assembly x86** (skill): Bootloaders and kernel work in indieOS. - **#frontend** (hub): 6 skills in Frontend. - **React** (skill): Default UI library for every web product I ship. - **Next.js** (skill): Full-stack framework of choice — App Router, RSC, edge-ready. - **TailwindCSS** (skill): Styling layer for every Next.js / React surface I build. - **SwiftUI** (skill): Declarative UI for native macOS and iOS apps. - **Flutter** (skill): Cross-platform mobile when Swift isn't the right tool. - **Tauri** (skill): Lightweight Rust-backed desktop shell — used for RustyExplorer, draw, dbquery, codehaven. - **#backend** (hub): 9 skills in Backend. - **Nest.js** (skill): Structured Node.js backend framework — used at scale for SynthAiLabs services. - **Flask** (skill): Lightweight Python web framework for quick AI inference endpoints. - **http (rust)** (skill): Low-level HTTP primitives in Rust for custom servers. - **Axum** (skill): Async Rust web framework powering the AI Inference Platform. - **Tokio (rust)** (skill): Async runtime underpinning every Rust service I write. - **Node.js** (skill): Backend runtime for TypeScript services, CLIs, and dev tooling. - **FastAPI** (skill): Async Python framework for ML serving and rapid API prototyping. - **Express** (skill): Minimal Node.js framework — used in MockServer's mock-engine. - **Prisma** (skill): Type-safe ORM for PostgreSQL-backed Next.js apps. - **#database** (hub): 8 skills in Database. - **PostgreSQL** (skill): Primary relational store for production apps. - **MongoDB** (skill): Document store for flexible-schema workloads. - **TimescaleDB** (skill): Time-series Postgres extension for metrics and event data. - **Milvus** (skill): Vector database for high-recall semantic search at scale. - **Redis** (skill): Cache, message queue, and L2 layer for the AI Inference Platform. - **MySQL** (skill): Used in legacy/contract setups where Postgres isn't the default. - **pgvector** (skill): Embedding storage and similarity search inside Postgres. - **OpenSearch** (skill): Powered low-latency profile retrieval for the WhatsApp AI agents at Tejas AI. - **#devops** (hub): 10 skills in DevOps & Cloud. - **Docker** (skill): Default packaging for every backend service I deploy. - **Kubernetes** (skill): Container orchestration for horizontally scaled inference workloads. - **AWS** (skill): Primary cloud — EC2, S3, RDS, Lambda for production deploys. - **Azure** (skill): Used selectively for client engagements requiring Microsoft stack. - **GCP** (skill): BigQuery, Cloud Run, and Vertex AI when the workload fits. - **ShaktiCloud** (skill): India-hosted GPU cloud — used at AuraX for diffusion inference. - **CI/CD** (skill): Build pipelines, automated testing, and zero-touch deploys. - **GitHub Actions** (skill): Workflow engine of choice — covers every repo I ship. - **Prometheus** (skill): Metrics scraping for the AI Inference Platform observability stack. - **Grafana** (skill): Dashboards on top of Prometheus for latency, cache, and model health. - **#ai** (hub): 14 skills in AI/ML/QM. - **PyTorch** (skill): Core deep-learning framework for VTO research and custom models. - **TensorFlow** (skill): Used selectively where the ecosystem fits (Keras, TFLite). - **scikit-learn** (skill): Classical ML for baselines and feature engineering. - **diffusers** (skill): HuggingFace library used heavily for Flux-based VTO pipelines. - **CUDA** (skill): Custom kernels and inference tuning on NVIDIA GPUs. - **Unsloth** (skill): Memory-efficient fine-tuning library — powers TuneSLM. - **Qiskit** (skill): IBM's quantum SDK — used for entanglement and QFT experiments. - **OpenQASM** (skill): Quantum assembly for circuit-level optimization in QM research. - **LLM** (skill): General large language model stack — OpenAI, Anthropic, Gemini, local Ollama. - **LoRA** (skill): Low-rank adapters — central to AuraX's C-AAC composition algorithm. - **Ollama** (skill): Local LLM runtime — used in localai, codehaven, and dbquery. - **Gemini API** (skill): Google's LLM API — used in dbquery and Syntaxia. - **OpenAPI** (skill): Spec-driven backend tooling — powers OpenCrow's server-side tool layer. - **Mem0** (skill): Persistent conversational memory layer built for the Tejas AI WhatsApp agents. ## Edges - ~/himanshu → #projects - ~/himanshu → #research - ~/himanshu → #experience - ~/himanshu → #skills - #skills → #sk-languages - #sk-languages → skill:TypeScript - #sk-languages → skill:Python - #sk-languages → skill:Rust - #sk-languages → skill:Swift - #sk-languages → skill:Dart - #sk-languages → skill:C/C++ - #sk-languages → skill:Go - #sk-languages → skill:Assembly x86 - #skills → #sk-frontend - #sk-frontend → skill:React - #sk-frontend → skill:Next.js - #sk-frontend → skill:TailwindCSS - #sk-frontend → skill:SwiftUI - #sk-frontend → skill:Flutter - #sk-frontend → skill:Tauri - #skills → #sk-backend - #sk-backend → skill:Nest.js - #sk-backend → skill:Flask - #sk-backend → skill:http (rust) - #sk-backend → skill:Axum - #sk-backend → skill:Tokio (rust) - #sk-backend → skill:Node.js - #sk-backend → skill:FastAPI - #sk-backend → skill:Express - #sk-backend → skill:Prisma - #skills → #sk-database - #sk-database → skill:PostgreSQL - #sk-database → skill:MongoDB - #sk-database → skill:TimescaleDB - #sk-database → skill:Milvus - #sk-database → skill:Redis - #sk-database → skill:MySQL - #sk-database → skill:pgvector - #sk-database → skill:OpenSearch - #skills → #sk-devops-cloud - #sk-devops-cloud → skill:Docker - #sk-devops-cloud → skill:Kubernetes - #sk-devops-cloud → skill:AWS - #sk-devops-cloud → skill:Azure - #sk-devops-cloud → skill:GCP - #sk-devops-cloud → skill:ShaktiCloud - #sk-devops-cloud → skill:CI/CD - #sk-devops-cloud → skill:GitHub Actions - #sk-devops-cloud → skill:Prometheus - #sk-devops-cloud → skill:Grafana - #skills → #sk-ai-ml-qm - #sk-ai-ml-qm → skill:PyTorch - #sk-ai-ml-qm → skill:TensorFlow - #sk-ai-ml-qm → skill:scikit-learn - #sk-ai-ml-qm → skill:diffusers - #sk-ai-ml-qm → skill:CUDA - #sk-ai-ml-qm → skill:Unsloth - #sk-ai-ml-qm → skill:Qiskit - #sk-ai-ml-qm → skill:OpenQASM - #sk-ai-ml-qm → skill:LLM - #sk-ai-ml-qm → skill:LoRA - #sk-ai-ml-qm → skill:Ollama - #sk-ai-ml-qm → skill:Gemini API - #sk-ai-ml-qm → skill:OpenAPI - #sk-ai-ml-qm → skill:Mem0 - #experience → co:tejas-ai - co:tejas-ai → skill:TypeScript - co:tejas-ai → skill:Go - co:tejas-ai → skill:Node.js - co:tejas-ai → skill:Nest.js - co:tejas-ai → skill:PostgreSQL - co:tejas-ai → skill:Redis - co:tejas-ai → skill:OpenSearch - co:tejas-ai → skill:Mem0 - co:tejas-ai → skill:LLM - co:tejas-ai → skill:Docker - co:tejas-ai → skill:AWS - #experience → co:aurax - co:aurax → skill:Python - co:aurax → skill:PyTorch - co:aurax → skill:diffusers - co:aurax → skill:LoRA - co:aurax → skill:CUDA - co:aurax → skill:LLM - co:aurax → skill:FastAPI - co:aurax → skill:PostgreSQL - co:aurax → skill:Redis - co:aurax → skill:Docker - co:aurax → skill:AWS - co:aurax → skill:ShaktiCloud - co:aurax → skill:Next.js - co:aurax → skill:Nest.js - co:aurax → res:c-aac - co:aurax → res:renu-virtual-try-on - #experience → co:synthailabs - co:synthailabs → skill:TypeScript - co:synthailabs → skill:Next.js - co:synthailabs → skill:Node.js - co:synthailabs → skill:Express - co:synthailabs → skill:MongoDB - co:synthailabs → skill:PostgreSQL - co:synthailabs → skill:Redis - co:synthailabs → skill:Docker - co:synthailabs → skill:CI/CD - co:synthailabs → skill:React - co:synthailabs → skill:TailwindCSS - co:synthailabs → res:moe-brain - #experience → co:tublian - co:tublian → skill:Python - co:tublian → skill:LLM - co:tublian → skill:Ollama - co:tublian → skill:FastAPI - co:tublian → skill:Docker - co:tublian → skill:TypeScript - co:tublian → skill:React - co:tublian → skill:Node.js - co:tublian → skill:GitHub Actions - #experience → co:stealth - co:stealth → skill:Python - co:stealth → skill:Qiskit - co:stealth → skill:OpenQASM - co:stealth → skill:C/C++ - #experience → co:excelerate - co:excelerate → skill:TypeScript - co:excelerate → skill:React - co:excelerate → skill:Next.js - #projects → #pr-ai - #pr-ai → proj:CCProxy - proj:CCProxy → skill:TypeScript - #pr-ai → proj:OpenCrow - proj:OpenCrow → skill:React - proj:OpenCrow → skill:TypeScript - proj:OpenCrow → skill:Node.js - proj:OpenCrow → skill:OpenAPI - proj:OpenCrow → skill:LLM - #projects → #pr-infra - #pr-infra → proj:AI Inference Platform - proj:AI Inference Platform → skill:Rust - proj:AI Inference Platform → skill:Axum - proj:AI Inference Platform → skill:Tokio (rust) - proj:AI Inference Platform → skill:Redis - proj:AI Inference Platform → skill:Prometheus - proj:AI Inference Platform → skill:Grafana - proj:AI Inference Platform → skill:Docker - #pr-ai → proj:TuneSLM - proj:TuneSLM → skill:Python - proj:TuneSLM → skill:Unsloth - proj:TuneSLM → skill:LLM - #projects → #pr-devtools - #pr-devtools → proj:MockServer - proj:MockServer → skill:TypeScript - proj:MockServer → skill:Next.js - proj:MockServer → skill:Express - #pr-devtools → proj:FastSearch - proj:FastSearch → skill:Rust - proj:FastSearch → skill:TypeScript - #pr-ai → proj:FallenAngel-SE/CY - proj:FallenAngel-SE/CY → skill:LLM - proj:FallenAngel-SE/CY → skill:Python - proj:FallenAngel-SE/CY → skill:TypeScript - #pr-devtools → proj:Dynamo-Prisma - proj:Dynamo-Prisma → skill:TypeScript - proj:Dynamo-Prisma → skill:Prisma - proj:Dynamo-Prisma → skill:Node.js - #projects → #pr-learning - #pr-learning → proj:ScratchAI-Algo - proj:ScratchAI-Algo → skill:Python - co:synthailabs → proj:ScratchAI-Algo - #pr-learning → proj:QM - Research - proj:QM - Research → skill:Python - proj:QM - Research → skill:Qiskit - proj:QM - Research → skill:OpenQASM - #projects → #pr-apps - #pr-apps → proj:RustyExplorer - proj:RustyExplorer → skill:Rust - proj:RustyExplorer → skill:TypeScript - proj:RustyExplorer → skill:Tauri - proj:RustyExplorer → skill:TailwindCSS - proj:RustyExplorer → skill:React - #pr-learning → proj:Linux Mini-Firewall - proj:Linux Mini-Firewall → skill:C/C++ - #pr-learning → proj:indieOS - proj:indieOS → skill:Assembly x86 - proj:indieOS → skill:C/C++ - #pr-devtools → proj:Codeforces Scraper - proj:Codeforces Scraper → skill:Python - #pr-devtools → proj:Log Ingestor - proj:Log Ingestor → skill:Rust - #pr-apps → proj:ZenStore (In-memory DB) - proj:ZenStore (In-memory DB) → skill:Rust - #pr-devtools → proj:frm (Fast Delete) - proj:frm (Fast Delete) → skill:Rust - #pr-apps → proj:roadmap.ai - proj:roadmap.ai → skill:Next.js - proj:roadmap.ai → skill:LLM - #pr-devtools → proj:biv - proj:biv → skill:Rust - #pr-learning → proj:ruspy - proj:ruspy → skill:Rust - #pr-apps → proj:PoW-Shield (DDoS Prevention) - proj:PoW-Shield (DDoS Prevention) → skill:C/C++ - proj:PoW-Shield (DDoS Prevention) → skill:Rust - proj:PoW-Shield (DDoS Prevention) → skill:TypeScript - #pr-devtools → proj:git-history - proj:git-history → skill:Rust - proj:git-history → skill:Python - #pr-apps → proj:localai - proj:localai → skill:Tauri - proj:localai → skill:Rust - proj:localai → skill:Next.js - proj:localai → skill:Ollama - #pr-apps → proj:draw (Excalidraw for Mac) - proj:draw (Excalidraw for Mac) → skill:Tauri - proj:draw (Excalidraw for Mac) → skill:Rust - proj:draw (Excalidraw for Mac) → skill:Next.js - #pr-apps → proj:Syntaxia (learncode) - proj:Syntaxia (learncode) → skill:Tauri - proj:Syntaxia (learncode) → skill:Rust - proj:Syntaxia (learncode) → skill:Next.js - proj:Syntaxia (learncode) → skill:Ollama - proj:Syntaxia (learncode) → skill:Gemini API - #pr-apps → proj:dbquery - proj:dbquery → skill:Tauri - proj:dbquery → skill:Rust - proj:dbquery → skill:Next.js - proj:dbquery → skill:Ollama - proj:dbquery → skill:Gemini API - #pr-apps → proj:codehaven.org - proj:codehaven.org → skill:Tauri - proj:codehaven.org → skill:Rust - proj:codehaven.org → skill:Next.js - proj:codehaven.org → skill:Ollama - #research → res:mmap-vs-read-apple-silicon - res:mmap-vs-read-apple-silicon → skill:C/C++ - res:mmap-vs-read-apple-silicon → skill:Rust - res:mmap-vs-read-apple-silicon → skill:PostgreSQL - #research → res:c-aac - res:c-aac → skill:LLM - res:c-aac → skill:diffusers - res:c-aac → skill:LoRA - #research → res:renu-virtual-try-on - res:renu-virtual-try-on → skill:diffusers - res:renu-virtual-try-on → skill:LoRA - #research → res:moe-brain - res:moe-brain → skill:LLM - #research → res:lattice-based-quantum-encryption - #research → res:gpt-neo - res:gpt-neo → skill:PyTorch - #research → res:custom-n-n-mnist - res:custom-n-n-mnist → skill:PyTorch --- Source: https://himanshuat.com/graph ======================================================================== # Fintracker A minimal, offline-first personal finance app built with Flutter. Track income, categorise spending and manage accounts, with everything stored locally on the device — no accounts, no cloud sync, no servers. Status: in verification on Android. ## Why it exists - **Privacy first.** Data never leaves the device. - **Smart analytics.** Spending visualised as charts. - **Local only.** No sign-up, no backend to trust. Privacy policy: https://himanshuat.com/app/fintracker/privacy --- Source: https://himanshuat.com/app/fintracker ======================================================================== # Fintracker — Privacy Policy Fintracker stores all financial data locally on your device. There are no user accounts, no cloud servers and no transmission of your data off the device. > This is a plain-text summary for convenience. The page below is the > authoritative version of the policy — read it there before relying on it. https://himanshuat.com/app/fintracker/privacy --- Source: https://himanshuat.com/app/fintracker/privacy ======================================================================== --- title: "A Beginner's Guide to Linux From Scratch" description: "Learn how to build a Linux system from scratch without ISO images." date: "February 16, 2021" url: "https://himanshuat.com/blogs/A-Beginnner-Guide-to-Linux-From-Scratch" --- # A Beginner's Guide to Linux From Scratch ## Objective 1. Gain a top-level understanding of how Linux works and its foundational principles. 2. Gather useful tips and tricks from someone who has already gone through the process. 3. Learn to install Linux on any hardware without using ISO images. For reference, I have used the following blogs, which provide some of the best open-source documentation available: [Linux From Scratch](http://www.linuxfromscratch.org/lfs/) ## My Host Machines During Compilation - **Round 1:** Dedicated host with base OS as Ubuntu - **Round 2:** Ubuntu VM installed on a base Ubuntu-operated laptop ## Let's Get Started ### Step 1: Creating a Dedicated Partition for LFS Compilation and Base Installation We will compile all the packages needed to build a Linux system and install our Linux OS in the same directory. The utility I used to make partitions is an app called Gparted, which is preinstalled in the Ubuntu flavor. Here's a quick overview of the partition setup: - **Root mount:** `/dev/sda1` - **Swap memory:** 8GB partition `/dev/sda5` - **Unallocated space:** 20GB for LFS compilation and installation (recommended to create a partition of about 30-35 GB). #### Commands for Partition Setup: ```bash export LFS=/mnt/lfs mkdir $LFS echo $LFS # should output /mnt/lfs mkdir $LFS/sources mkdir $LFS/tools ln -sv $LFS/tools / ``` ### Step 2: Download LFS Binaries from Official LFS Repos ```bash cd $LFS/sources wget http://www.linuxfromscratch.org/lfs/downloads/8.4/wget-list wget http://www.linuxfromscratch.org/lfs/downloads/8.4/md5sums pushd $LFS/sources md5sum -c md5sums ``` The above command should return all packages checked with "OK" as a result. ### Step 3: Adding LFS User and Setting Up the Environment ```bash groupadd lfs useradd -s /bin/bash -g lfs -m -k /dev/null lfs passwd lfs # Give a password for lfs user login chown -v lfs $LFS/tools chown -v lfs $LFS/sources ``` #### Environment Setup: ```bash su - lfs cat > ~/.bash_profile << "EOF" exec env -i HOME=$HOME TERM=$TERM PS1='\u:\w\$ ' /bin/bash EOF cat > ~/.bashrc << "EOF" set +h umask 022 LFS=/mnt/lfs LC_ALL=POSIX LFS_TGT=$(uname -m)-lfs-linux-gnu PATH=/tools/bin:/bin:/usr/bin export LFS LC_ALL LFS_TGT PATH EOF source ~/.bash_profile echo $LFS # should return /mnt/lfs ``` ### Step 4: General Steps for the Compilation Process For compiling all the packages/binaries, follow these steps: ```bash cd $LFS/sources tar -xvf cd ./configure && make make check # (optional for packages with test cases) make install rm -rf ``` Compile all packages in the exact order given below and follow every step mentioned in the LFS book: [Linux From Scratch - Book](http://www.linuxfromscratch.org/lfs/downloads/8.4/LFS-BOOK-8.4-NOCHUNKS.html) #### Example Packages to Compile: - Binutils-2.32 (Pass 1) - GCC-8.2.0 (Pass 1) - Linux-4.20.12 API Headers - Glibc-2.29 - Libstdc++ from GCC-8.2.0 - ... (and many more as listed in the LFS book) ### Step 5: Continued Compilation Continue compiling the next set of packages as mentioned in the LFS book. Ensure you follow the exact sequence and steps. ### Step 6: Udev, Hostname, IP Configuration, and Final Steps for Booting Linux After compiling the Linux kernel, it's essential to set up network scripts correctly to maintain connectivity. #### Final Command for Boot Configuration: ```bash cat > /boot/grub/grub.cfg << "EOF" # Begin /boot/grub/grub.cfg set default=0 set timeout=5 insmod ext2 set root=(hd0,2) menuentry "GNU/Linux, Linux 4.20.12-lfs-8.4" { linux /boot/vmlinuz-4.20.12-lfs-8.4 root=/dev/sda2 ro } EOF ``` ### General Tips for the Entire Journey - Don’t deviate from the book at any point. - Write sensible network configuration scripts if your host system is not accessible geographically. - Be patient and maintain a positive attitude throughout the process. ## Conclusion Thanks for reading my blog. Keep coding, keep smiling! --- Source: https://himanshuat.com/blogs/A-Beginnner-Guide-to-Linux-From-Scratch ======================================================================== --- title: "Beyond Naive RAG: Advanced Chunking and Embedding Strategies for Superior Retrieval" description: "A fixed-size chunker splits the sentence that answers the query across two chunks, and neither half retrieves. The fact is in your corpus and the system says it doesn't know. Chunking strategy and embedding choice are the two levers that fix that, and Hit Rate and MRR are how you prove it." date: "October 28, 2025" url: "https://himanshuat.com/blogs/advanced-rag-chunking-embedding-strategies" --- # Beyond Naive RAG: Advanced Chunking and Embedding Strategies for Superior Retrieval A fixed-size chunker cuts a technical manual into 500-character slices. It doesn't look at where sentences end, or where paragraphs end, or where chapters end. So the sentence that answers the question lands half in one chunk and half in the next. Neither half retrieves cleanly. The fact is sitting in your corpus, and your system tells the user it doesn't know. > Most RAG failures are ingestion failures. They are decided before a single query runs, by where you cut the text and what you embed it with. Two things get broken at that boundary. Important information gets split across it, and unrelated text gets bundled in on the other side as noise. Both are permanent by the time a query arrives. We don't store facts in isolation. We organize them in context, so recalling one thing pulls in related ideas and surrounding detail. A naive RAG system hands back disjointed fragments instead. This matters most for MoE architectures, where each specialized expert needs focused, relevant context to do its job. Weak retrieval starves an expert of what it needs, and its decisions get fuzzy. ### Chapter 0 — A chunk is the unit of retrieval Whatever the retriever hands the model, it hands over whole. It cannot return two thirds of a chunk, and it cannot stitch the missing third back on from the chunk next door. That's the whole reason chunk boundaries carry so much weight. **The boundary decides what can ever be retrieved together.** Everything below is about putting that boundary somewhere defensible. --- ### Step 1 — Align the cut to a natural break Fixed-size chunking is the baseline, limitations and all. Don't start by replacing it. Start by fixing where it lands. Instead of a flat character count, snap the boundary back to the nearest sentence or paragraph end that still falls inside the window. The chunk comes out slightly shorter and stops mid-thought far less often. ```typescript type TextChunk = { content: string; metadata?: Record; }; /** * Chunks text by a fixed character size with a specified overlap. * Ensures chunks attempt to end/start on natural breaks (sentences) within overlap. * @param text The input document text. * @param chunkSize The maximum size of each chunk. * @param overlapSize The size of the overlap between chunks. * @returns An array of TextChunk objects. */ function fixedSizeChunker(text: string, chunkSize: number, overlapSize: number = 0): TextChunk[] { if (overlapSize >= chunkSize) { throw new Error("Overlap size must be less than chunk size."); } const chunks: TextChunk[] = []; let currentPos = 0; while (currentPos < text.length) { let endPos = Math.min(currentPos + chunkSize, text.length); let startPos = currentPos; // Adjust endPos to not cut sentences in half if possible, within the chunk boundary if (endPos < text.length) { const sentenceEnd = text.lastIndexOf('.', endPos - 1) + 1; const paragraphEnd = text.lastIndexOf('\n\n', endPos - 1) + 2; const lastRelevantBreak = Math.max(sentenceEnd, paragraphEnd); if (lastRelevantBreak > startPos && lastRelevantBreak < endPos) { endPos = lastRelevantBreak; } } const chunkContent = text.substring(startPos, endPos).trim(); if (chunkContent) { chunks.push({ content: chunkContent }); } // Move currentPos forward, accounting for overlap currentPos = endPos - overlapSize; if (currentPos <= startPos) { // Handle cases where overlap pushes currentPos back too far currentPos = startPos + chunkSize; // Force progress if chunk size is small or overlap large } } return chunks; } // Example usage // const doc = "This is the first sentence. This is the second sentence. And here is the third sentence. A new paragraph starts here. With more content."; // const fixedChunks = fixedSizeChunker(doc, 100, 20); // console.log("Fixed Chunks:", fixedChunks); ``` That's better than `text.substring(i, i + size)`, and it is still a heuristic. It has no idea what it's cutting through. A period is a weak proxy for a topic ending, and the function's only defence against a pathological overlap setting is the forced-progress branch at the bottom, which is a trapdoor rather than a fix. **The tell:** print ten retrieved chunks. If any of them start mid-sentence, your boundaries are being decided by a character counter. --- ### Step 2 — Cut where the topic shifts, not where the counter runs out This is where the real gains are. Semantic chunking keeps related ideas together by finding the break instead of scheduling it. The method I reach for: 1. Split the text into elementary units, usually sentences. 2. Embed each unit with a lightweight model. 3. Find the semantic breaks, the points where the topic shifts. A large drop in cosine similarity between adjacent sentence embeddings is a good signal. 4. Group the sentences between breaks into larger, coherent chunks. The shape of it is a loop with one decision in it. Every sentence either joins the chunk being built or ends it. ```mermaid flowchart TD A[split sentences] --> B[embed each] B --> C[append to chunk] C --> D{similarity to next} D -->|above threshold| C D -->|below threshold| E[emit and reset] E --> C ``` The implementation follows that diagram directly. The threshold and the minimum chunk size are the two knobs, defaulting to `0.7` and three sentences. ```typescript // Assuming a simple sentence splitter (e.g., from Agno or a custom regex) // For simplicity, I'll use a basic split for this example, but a robust solution // needs a proper NLP sentence tokenizer. function splitIntoSentences(text: string): string[] { // A simple regex split, not production-ready for all languages/edge cases return text.match(/[^.!?]+[.!?]+/g) || []; } // Mock embedding function for demonstration // In a real scenario, this would call an API or a local model. async function getSentenceEmbedding(sentence: string): Promise { // Simulate an embedding call const encoder = new TextEncoder(); const data = encoder.encode(sentence); let hash = 0; for (let i = 0; i < data.length; i++) { hash = (hash << 5) - hash + data[i]; hash |= 0; // Convert to 32bit integer } return [hash % 1000 / 1000, (hash * 2) % 1000 / 1000, (hash * 3) % 1000 / 1000]; // Dummy 3D embedding } function cosineSimilarity(vecA: number[], vecB: number[]): number { let dotProduct = 0; let magnitudeA = 0; let magnitudeB = 0; for (let i = 0; i < vecA.length; i++) { dotProduct += vecA[i] * vecB[i]; magnitudeA += vecA[i] * vecA[i]; magnitudeB += vecB[i] * vecB[i]; } magnitudeA = Math.sqrt(magnitudeA); magnitudeB = Math.sqrt(magnitudeB); if (magnitudeA === 0 || magnitudeB === 0) return 0; return dotProduct / (magnitudeA * magnitudeB); } /** * Chunks text semantically by detecting significant topic shifts. * @param text The input document text. * @param similarityThreshold Threshold below which a similarity score indicates a topic shift (e.g., 0.6-0.7). * @param minChunkSize Minimum number of sentences per chunk. * @returns An array of TextChunk objects. */ async function semanticChunker( text: string, similarityThreshold: number = 0.7, minChunkSize: number = 3 ): Promise { const sentences = splitIntoSentences(text.trim()); if (sentences.length === 0) return []; const sentenceEmbeddings: number[][] = await Promise.all( sentences.map(s => getSentenceEmbedding(s)) ); const chunks: TextChunk[] = []; let currentChunkSentences: string[] = []; let currentChunkEmbeddings: number[][] = []; for (let i = 0; i < sentences.length; i++) { currentChunkSentences.push(sentences[i]); currentChunkEmbeddings.push(sentenceEmbeddings[i]); // If we have enough sentences, check for a semantic break if (currentChunkSentences.length >= minChunkSize && i < sentences.length - 1) { const currentAvgEmbedding = currentChunkEmbeddings.reduce( (acc, val) => acc.map((v, idx) => v + val[idx]), Array(sentenceEmbeddings[0].length).fill(0) ).map(v => v / currentChunkEmbeddings.length); const nextSentenceEmbedding = sentenceEmbeddings[i + 1]; const similarityToNext = cosineSimilarity(currentAvgEmbedding, nextSentenceEmbedding); if (similarityToNext < similarityThreshold) { // Semantic break detected chunks.push({ content: currentChunkSentences.join(' ').trim() }); currentChunkSentences = []; currentChunkEmbeddings = []; } } } // Add any remaining sentences as the last chunk if (currentChunkSentences.length > 0) { chunks.push({ content: currentChunkSentences.join(' ').trim() }); } return chunks; } // Example usage (will need an actual embedding model) // const complexDoc = "Quantum entanglement is a phenomenon where two particles become linked. Their states are interdependent. This has profound implications for quantum computing. On a completely different note, the stock market closed higher today. Technology stocks saw significant gains."; // (async () => { // const semanticChunks = await semanticChunker(complexDoc, 0.7); // console.log("Semantic Chunks:", semanticChunks); // })(); ``` The comparison is against a running average of the chunk so far, not against the previous sentence alone. That's deliberate: it makes a single odd sentence less likely to trigger a spurious break. It costs more compute than counting characters, because every sentence gets embedded at ingest time before a single chunk exists. For an MoE-driven agent, where each expert only reasons well on clean context, that trade is worth making. **The tell:** print the boundaries for one document you know well. If you can't name each chunk's topic in a phrase, the splitter isn't finding topics. --- ### Step 3 — Pick the embedding model against a constraint you can name The embedding model turns text into vectors that capture meaning. A weak one produces vectors that can't separate relevant text from irrelevant text, and at that point even the best chunking won't save you. The choice is a constraint problem, not a ranking problem. | model | origin | why you'd pick it | |---|---|---| | `bge-small-en-v1.5` | BAAI, open source | compact, punches above its size; on-device or resource-constrained setups | | `E5-large-v2` | Microsoft, open source | larger, needs more resources, gives better embedding quality | | `all-MiniLM-L6-v2` | open source | smaller and faster, with a fair speed-to-quality balance | | `text-embedding-ada-002` | OpenAI, proprietary | the previous standard, widely used for quality and easy access | | `text-embedding-3-small` / `-3-large` | OpenAI, proprietary | newer generation, better performance, output dimensions configurable against cost | The `text-embedding-3` pair usually tops the quality rankings for what they cost. When you need local processing or tight latency, fine-tuned open-source models like `bge-small-en-v1.5` or `E5-large-v2` hold their own. **The tell:** if you can't say which constraint picked your model — latency, hardware, cost, or a measured score on your own data — you picked it off a leaderboard. --- ### Step 4 — Score the change, don't feel it To compare chunking and embedding choices you need real metrics, not an anecdotal "it feels better." Start with a small, representative set of documents and query-relevance pairs, where each query has its *ground truth* relevant chunks marked in the source document. That labelling is the expensive part and there is no way around it. Two numbers do the work. **Hit Rate**, the proportion of queries where at least one of the top `k` retrieved chunks is relevant: $$ HitRate = \frac{1}{N} \sum_{i=1}^{N} \mathbb{I}(\text{relevant chunk in top k for query } i) $$ And **Mean Reciprocal Rank**, the average of the reciprocal ranks of the first relevant document, where $rank_q$ is the rank of the first relevant document for query $q$: $$ MRR = \frac{1}{|Q|} \sum_{q=1}^{|Q|} \frac{1}{rank_q} $$ Hit Rate tells you whether the right chunk came back at all. MRR tells you how far down the list it was. A change that improves one and flattens the other is still worth knowing about. The harness below embeds every chunk once, retrieves by cosine similarity, and reports both numbers for whatever chunking strategy produced the chunks. ```typescript // Assume 'documents' is an array of strings, 'queries' is an array of strings, // and 'groundTruth' is a Map> where key=query, value=set of original relevant chunk texts. interface Retriever { retrieve(query: string, k: number): Promise; } // Dummy Retriever implementation for demonstration class DummyRetriever implements Retriever { private chunks: TextChunk[]; // All document chunks for a given strategy private embeddings: Map; // Pre-computed chunk embeddings private embeddingFunction: (text: string) => Promise; constructor( chunks: TextChunk[], embeddingFunction: (text: string) => Promise ) { this.chunks = chunks; this.embeddingFunction = embeddingFunction; this.embeddings = new Map(); } async initialize() { for (const chunk of this.chunks) { this.embeddings.set(chunk.content, await this.embeddingFunction(chunk.content)); } } async retrieve(query: string, k: number): Promise { const queryEmbedding = await this.embeddingFunction(query); const similarities: { chunk: TextChunk; score: number }[] = []; for (const chunk of this.chunks) { const chunkEmbedding = this.embeddings.get(chunk.content); if (chunkEmbedding) { similarities.push({ chunk: chunk, score: cosineSimilarity(queryEmbedding, chunkEmbedding) }); } } // Sort by similarity and return top k return similarities .sort((a, b) => b.score - a.score) .slice(0, k) .map(s => s.chunk); } } async function runEvaluation( retriever: Retriever, queries: string[], groundTruth: Map>, k: number = 5 ): Promise<{ hitRate: number; mrr: number }> { let totalHit = 0; let totalReciprocalRank = 0; for (const query of queries) { const retrievedChunks = await retriever.retrieve(query, k); const relevantChunks = groundTruth.get(query) || new Set(); let hit = false; let rankOfFirstRelevant = 0; for (let i = 0; i < retrievedChunks.length; i++) { const retrievedContent = retrievedChunks[i].content; // Check if any part of the retrieved content overlaps significantly with a ground truth chunk // (This requires a more robust 'isRelevant' check than exact string match in real-world) const isRelevant = Array.from(relevantChunks).some(gt => retrievedContent.includes(gt) || gt.includes(retrievedContent)); if (isRelevant) { hit = true; if (rankOfFirstRelevant === 0) { rankOfFirstRelevant = i + 1; // 1-based rank } } } if (hit) { totalHit++; } if (rankOfFirstRelevant > 0) { totalReciprocalRank += 1 / rankOfFirstRelevant; } } const hitRate = totalHit / queries.length; const mrr = totalReciprocalRank / queries.length; return { hitRate, mrr }; } // Example usage with a real embedding function (e.g., calling OpenAI API) // class OpenAIBaseEmbedding implements (text: string) => Promise { // private apiKey: string; // constructor(apiKey: string) { this.apiKey = apiKey; } // async embed(text: string): Promise { // const response = await fetch("https://api.openai.com/v1/embeddings", { // method: "POST", // headers: { // "Content-Type": "application/json", // "Authorization": `Bearer ${this.apiKey}`, // }, // body: JSON.stringify({ // input: text, // model: "text-embedding-3-small", // }), // }); // const data = await response.json(); // return data.data[0].embedding; // } // } // (async () => { // const documentText = ` // Quantum computing is a rapidly emerging field that harnesses the principles of quantum mechanics to solve complex computational problems. Unlike classical computers which use bits, quantum computers use qubits, which can exist in multiple states simultaneously due to superposition. This allows quantum computers to process vast amounts of information in parallel. // One of the most promising applications of quantum computing is in drug discovery and materials science. Simulating molecular interactions with classical computers is extremely difficult, but quantum computers could revolutionize this area by accurately modeling complex chemical reactions. // Another significant application is in cryptography. Shor's algorithm, for example, can efficiently factor large numbers, posing a threat to many current encryption methods like RSA. However, quantum-safe cryptography is also an active area of research. // The development of quantum hardware is still in its early stages, with challenges in maintaining qubit coherence and scalability. Various technologies, including superconducting circuits, trapped ions, and photonic systems, are being explored to build stable quantum processors. // Beyond quantum computing, artificial intelligence continues to advance at an unprecedented pace. Large Language Models (LLMs) are transforming how we interact with information, enabling natural language understanding and generation. The integration of LLMs with specialized knowledge bases via RAG is crucial for building robust AI systems. My personal research focuses on combining these with MoE architectures for more efficient and intelligent agents, mimicking human cognitive processes. // `; // const queries = [ // "What are the applications of quantum computing?", // "How do quantum computers differ from classical computers?", // "What are the challenges in quantum hardware development?", // "What is my research focus?", // A query specifically for the last paragraph // ]; // const groundTruth = new Map>([ // ["What are the applications of quantum computing?", new Set(["One of the most promising applications of quantum computing is in drug discovery and materials science.", "Another significant application is in cryptography."])], // ["How do quantum computers differ from classical computers?", new Set(["Unlike classical computers which use bits, quantum computers use qubits, which can exist in multiple states simultaneously due to superposition."])], // ["What are the challenges in quantum hardware development?", new Set(["The development of quantum hardware is still in its early stages, with challenges in maintaining qubit coherence and scalability."])], // ["What is my research focus?", new Set(["My personal research focuses on combining these with MoE architectures for more efficient and intelligent agents, mimicking human cognitive processes."])] // ]); // // === Run with Fixed-Size Chunks === // const fixedChunks = fixedSizeChunker(documentText, 300, 50); // const openAIEmbedder = new OpenAIBaseEmbedding("YOUR_OPENAI_API_KEY"); // Replace with actual API key // const fixedRetriever = new DummyRetriever(fixedChunks, openAIEmbedder.embed); // await fixedRetriever.initialize(); // const fixedResults = await runEvaluation(fixedRetriever, queries, groundTruth); // console.log("Fixed-Size Chunking Results:", fixedResults); // // === Run with Semantic Chunks === // const semanticChunks = await semanticChunker(documentText, 0.7, 3); // Re-run init with actual embedder // const semanticRetriever = new DummyRetriever(semanticChunks, openAIEmbedder.embed); // await semanticRetriever.initialize(); // const semanticResults = await runEvaluation(semanticRetriever, queries, groundTruth); // console.log("Semantic Chunking Results:", semanticResults); // })(); ``` The commented block at the bottom is the actual experiment: same document, same embedder, same queries, two chunkers. Only the chunking strategy varies, which is the only way the difference in Hit Rate and MRR means anything. *Both `DummyRetriever` and the `getSentenceEmbedding` used by `semanticChunker` are simplified. For a real benchmark, `getSentenceEmbedding` would call the actual embedding model under evaluation, and `DummyRetriever` would be initialized with chunks from a specific chunking strategy.* Across the benchmarks, semantic chunking beat naive fixed-size chunking, and the gap widened on documents that spanned several topics or carried fine detail. Fixed-size with smart overlap beat arbitrary splits, and it still lost the coherence that nuanced retrieval depends on. I'm not putting a single headline number on that gap here, because a number measured on my corpus tells you nothing about yours. Run the harness on your documents. **The tell:** if your evidence for a chunking change is that answers "feel better", you don't have evidence, you have a mood. --- ### Five things to build once the harness runs Once you can score a retrieval change, each of these becomes a measurable experiment rather than an opinion. - **Hierarchical RAG.** Embed at several granularities — document, section, paragraph, sentence — and retrieve based on how complex the query is. - **Graph RAG.** Model knowledge as a graph so you capture relationships between concepts instead of flat, linear text. - **Hybrid retrieval.** Combine dense vector search with sparse keyword search (BM25). - **Query expansion and rewriting.** Use an LLM to rephrase or expand a query before retrieval. - **Self-correction.** Let the LLM judge the retrieved chunks and ask for something more specific when they fall short. --- ### What breaks it **The sentence splitter poisons everything downstream.** `splitIntoSentences` is a regex over `.!?`. It is not production-ready across languages or edge cases, and abbreviations, decimals, and quoted dialogue all break it. Every embedding, every similarity score, and every boundary in Step 2 inherits that error. A proper NLP tokenizer is the cheapest reliability upgrade in the pipeline. **The forced-progress branch silently drops text.** In `fixedSizeChunker`, when the sentence-break snap pulls `endPos` back to within `overlapSize` of the start, `currentPos = endPos - overlapSize` fails to advance and the fallback jumps to `startPos + chunkSize` instead. That lands past `endPos`, and the text in between is never emitted. It doesn't throw. You find out when a retrieval that should have worked doesn't. **The benchmark grades itself with a substring match.** `isRelevant` checks whether the retrieved content contains a ground truth chunk or vice versa. Trim a chunk boundary by one character and a genuine hit reads as a miss; return a whole page and it reads as a hit. A benchmark this permissive will happily report an improvement that is really a change in chunk length. **The threshold is a free parameter with no principled default.** `0.7` is a starting point, not a finding. Set it too high and you shatter a coherent section; too low and you merge two topics into one chunk that retrieves for neither. `minChunkSize` compounds this by forbidding a break before three sentences, so a genuine topic shift on sentence two waits until sentence three. --- ### When semantic chunking is the wrong choice The machinery costs an embedding call per sentence at ingest, plus the labelling effort behind the ground truth set. Several situations don't repay that. - **Your documents are already the right size.** FAQ entries, support tickets, function docstrings. One unit is one chunk, there is no boundary to get wrong, and a splitter can only damage them. - **Each document covers one topic.** The gap widens on documents that span several topics. The corollary is that on a single-topic document there is no topic shift for the detector to find, and it spends compute confirming that. - **You have no ground truth yet.** Without query-relevance pairs you cannot tell whether the change helped. Label a small set first. Switching strategies before you can score them means trading one unmeasured system for another. - **The bottleneck is somewhere else.** If a weak embedding model can't separate relevant from irrelevant text, better boundaries don't rescue it. Fix the embedding first, then come back to the cuts. --- ### The shift Retrieval quality gets decided at ingest time, by two decisions most implementations make on autopilot: where the text is cut, and what turns it into vectors. Neither is recoverable at query time. Understanding how your data is structured before you embed it pays for itself. That's the piece I keep coming back to for MoE work, where every expert reasons in a narrow domain and needs precise context to do it. Feed it noisy, fragmented retrieval and the whole system gets slower and less accurate. - **Cut on meaning, not on a character counter.** A period is a proxy for a topic ending, and a bad one. - **Pick the embedding model against a named constraint.** Latency, hardware, cost, or a score you measured yourself. - **Score before and after, on your own corpus.** Hit Rate says whether it came back, MRR says how far down. Label twenty queries against your own documents this week and run both chunkers against them. That number is the only one in this post that applies to you. --- Source: https://himanshuat.com/blogs/advanced-rag-chunking-embedding-strategies ======================================================================== --- title: "From Solo Agent to Team Player: Architecting Multi-Agent Systems with LangGraph" description: "One model asked to research, draft, review, and revise in a single pass does all four badly. Splitting the job across specialized agents and wiring them into a state machine graph makes every handoff explicit, and turns the supervisor into the only prompt that really matters." date: "March 01, 2024" url: "https://himanshuat.com/blogs/architecting-multi-agent-systems-with-langgraph" --- # From Solo Agent to Team Player: Architecting Multi-Agent Systems with LangGraph Ask one model to research a topic, draft a summary, critique the draft, and revise it, all in the same pass. It will do all four. It will do all four badly. It hallucinates, drifts out of coherence across the steps, and has no real specialization in any of the roles it's playing. Forcing one model to act as researcher, writer, and editor at once is like asking one person to ace every Olympic event. Most of the push toward AGI fixates on scaling single, monolithic models. But the human brain isn't one colossal neural network. It's a distributed network of specialized modules, each an expert in its domain, wired together through complex pathways. That MoE (Mixture of Experts) architecture, where different parts of the brain activate for different tasks, is my research north star. > My bet on intelligent systems is smarter *architectures*, not larger models. That leads straight to multi-agent systems, a first crude approximation of distributed intelligence. We're not chaining prompts anymore. We're building teams. ### Chapter 0 — What a state machine graph is Three parts, and you need all three before you write an agent. **Nodes** are the agents, or the functions they perform. **Edges** are the flow of control between nodes, often conditional. **State** is a shared, mutable context that every node can read and update, the single source of truth for the whole system. The difference from a chain is the edge. A chain runs A to B to C and stops. A graph can route control back, which is what lets a reviewer send work to a reviser and the reviser send it back for another look. Here's what that buys you over a single pass: | Concern | One model, one pass | Supervisor graph | |---|---|---| | Specialization | One prompt covering every role | One prompt per agent, and its own model and tools if you want | | Iteration | Whatever comes out is the output | Review and revise cycles until the supervisor stops them | | Parallelism | None | Logical delegation, though a synchronous setup still runs sequentially | | Decomposition | The whole problem in one context | Sub-problems, each with a dedicated expert | ### Why LangGraph, and not LangChain A word on frameworks, since I'm about to use one. I'm a competitive programmer; my portfolio prioritizes raw APIs, performance, and minimal abstraction. Heavy frameworks, especially ones that bury the underlying logic under layers of wrappers, are a hard pass. LangChain, for the most part, feels like a complex abstraction searching for a problem. Defining cyclical *state machines* is the narrow case where I'll admit a DSL beats writing raw FSM logic from scratch. LangGraph is a tool for graph definition, not an all-encompassing LLM orchestration layer, and for modeling a dynamic collaborative workflow it fits. In production I'd probably port the core graph logic to a custom TypeScript implementation for type safety, performance, and control. For prototyping and showing the architecture, it's a pragmatic pick. ### The job — a research team that ships a technical summary Given a technical query, produce a concise, well-researched technical summary. That decomposes into six moves: 1. **Understand the request:** clarify the user's intent. 2. **Research:** gather relevant information. 3. **Draft:** synthesize research into a coherent summary. 4. **Review:** evaluate the draft for accuracy, clarity, and completeness. 5. **Refine:** incorporate feedback and improve the draft. 6. **Finalize:** deliver the polished summary. Which maps directly onto a supervisor agent coordinating specialized workers. --- ### Step 1 — Define the shared state before you write an agent This is the part people skip, and skipping it is why their agents hallucinate context or redo work another agent already did. A well-defined state is the blackboard everyone writes on. Use a `TypedDict` for explicit typing, even in Python. ```python from typing import TypedDict, List, Optional from langchain_core.messages import BaseMessage class AgentState(TypedDict): """ Represents the state of our multi-agent system. Shared context accessible and modifiable by all agents. """ query: str # The initial user query/topic research_results: List[str] # Accumulated research findings draft: Optional[str] # Current draft of the summary review_comments: List[str] # Feedback from the reviewer iterations: int # Number of review-refine cycles max_iterations: int # Maximum allowed review-refine cycles next_action: Optional[str] # Supervisor's decision on the next agent to invoke messages: List[BaseMessage] # For conversational history if needed (e.g., LangGraph's default) ``` Note what's in there that isn't content: `iterations`, `max_iterations`, and `next_action`. Control flow lives in the state alongside the work, which is what makes the routing decision inspectable instead of buried in a call stack. **The tell:** if two agents need to pass something to each other and it isn't a key in this dict, they're passing it through prompt text, and nothing outside the prompt can see it. --- ### Step 2 — Make each agent a function over state Every agent takes `AgentState` and returns an updated slice of it. The intelligence comes from an LLM call; the logic and the state manipulation stay pure Python. The raw `openai` client keeps the model call one function deep instead of behind a wrapper. ```python import os from openai import OpenAI from typing import Callable, Literal from langchain_core.messages import HumanMessage, SystemMessage, FunctionMessage from langgraph.graph import StateGraph, END # --- LLM Client Setup --- # Use raw OpenAI client for direct API interaction, avoiding LangChain's LLM wrappers. client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY")) def call_llm(prompt: str, model: str = "gpt-4o", temperature: float = 0.7) -> str: """Helper to make direct LLM calls.""" response = client.chat.completions.create( model=model, messages=[{"role": "user", "content": prompt}], temperature=temperature, ) return response.choices[0].message.content # --- Agent Definitions --- class Agents: """Contains our specialized agent functions.""" def __init__(self, llm_model: str = "gpt-4o"): self.llm_model = llm_model def supervisor_agent(self, state: AgentState) -> AgentState: """ The orchestrator. Decides which worker agent to invoke next or if the task is complete. Returns a state dict with 'next_action'. """ current_draft_status = "no draft yet" if state.get("draft"): current_draft_status = "draft exists" review_status = "no review comments" if state.get("review_comments"): review_status = f"{len(state['review_comments'])} review comments exist" supervisor_prompt = f""" You are the Supervisor. Your goal is to manage a research and content generation team. The current query is: "{state['query']}" Current Status: - Research Results: {len(state['research_results'])} pieces of information gathered. - Draft Status: {current_draft_status} - Review Status: {review_status} - Iterations: {state['iterations']}/{state['max_iterations']} Based on the current state, decide the next logical step. Choose one of the following actions: 'research', 'draft', 'review', 'revise', 'finish'. Rules: 1. If no research has been done, or existing research is insufficient, choose 'research'. 2. If research is sufficient and no draft exists, choose 'draft'. 3. If a draft exists and has not been reviewed (or needs further review), choose 'review'. 4. If review comments exist and iterations < max_iterations, choose 'revise'. 5. If a draft has been reviewed, all comments addressed (or max_iterations reached), choose 'finish'. Your output MUST be a single word, one of: research, draft, review, revise, finish. """ print(f"--- Supervisor Thinking ---") action = call_llm(supervisor_prompt, model=self.llm_model, temperature=0.1).strip().lower() print(f"Supervisor decided: {action}") return {"next_action": action} def researcher_agent(self, state: AgentState) -> AgentState: """ Gathers information based on the query. (For this example, we'll simulate web search with an LLM call.) """ print(f"--- Researcher Working ---") research_prompt = f""" You are a highly skilled researcher. Find 3-5 key facts or pieces of information about the following topic: "{state['query']}". Focus on technical aspects and provide concise bullet points. If previous research results exist, try to expand on them or find complementary information. Previous results: {state['research_results']} """ new_research = call_llm(research_prompt, model=self.llm_model).strip() updated_results = state.get("research_results", []) + [new_research] print(f"Researcher found: {new_research[:100]}...") return {"research_results": updated_results} def writer_agent(self, state: AgentState) -> AgentState: """ Drafts the technical summary based on research results. """ print(f"--- Writer Working ---") research_summary = "\n".join(state["research_results"]) writer_prompt = f""" You are a technical writer. Draft a concise, informative summary (approx. 200-300 words) on the topic: "{state['query']}". Use the following research results: --- {research_summary} --- Focus on clarity, accuracy, and technical detail appropriate for a developer audience. """ draft = call_llm(writer_prompt, model=self.llm_model).strip() print(f"Writer drafted: {draft[:100]}...") return {"draft": draft} def reviewer_agent(self, state: AgentState) -> AgentState: """ Reviews the draft and provides constructive criticism. """ print(f"--- Reviewer Working ---") draft = state["draft"] reviewer_prompt = f""" You are a critical technical reviewer. Evaluate the following draft for accuracy, clarity, completeness, and adherence to the topic "{state['query']}". Provide constructive feedback, suggesting specific improvements. If the draft is perfect, state "No specific changes needed." Draft: --- {draft} --- """ comments = call_llm(reviewer_prompt, model=self.llm_model).strip() updated_comments = state.get("review_comments", []) + [comments] print(f"Reviewer commented: {comments[:100]}...") return {"review_comments": updated_comments} def reviser_agent(self, state: AgentState) -> AgentState: """ Revises the draft based on reviewer comments. Increments the iteration count. """ print(f"--- Reviser Working ---") original_draft = state["draft"] review_feedback = "\n".join(state["review_comments"]) reviser_prompt = f""" You are a meticulous reviser. Improve the following draft based on the provided review feedback. Ensure all comments are addressed while maintaining the original intent and technical accuracy for the topic "{state['query']}". Original Draft: --- {original_draft} --- Review Feedback: --- {review_feedback} --- Provide the revised draft. """ revised_draft = call_llm(reviser_prompt, model=self.llm_model).strip() # Clear review comments for the next cycle and increment iteration new_iterations = state.get("iterations", 0) + 1 print(f"Reviser revised. Iteration: {new_iterations}") return {"draft": revised_draft, "review_comments": [], "iterations": new_iterations} ``` Two details carry more weight than the rest. The supervisor runs at `temperature=0.1` while the workers run at the default, because the supervisor isn't writing, it's choosing. And its prompt ends with a hard output contract: a single word from a closed set of five. That line is what makes the next step possible. **The tell:** every agent returns a partial state dict and nothing else. If one of yours returns a formatted string for another agent to parse, the state has stopped being the source of truth. --- ### Step 3 — Route every worker back through the supervisor The topology is a star with the supervisor at the center. Entry goes to the supervisor, a conditional edge fans out to whichever worker its decision names, and each worker's edge points straight back. ```mermaid flowchart TD S{supervisor} S -->|research| R[researcher] S -->|draft| W[writer] S -->|review| V[reviewer] S -->|revise| X[reviser] S -->|finish| E([END]) R --> S W --> S V --> S X --> S ``` That single cycle back to the center is the whole design. No worker ever decides what happens next. ```python # Initialize our agents team_agents = Agents() # Build the graph workflow = StateGraph(AgentState) # Add nodes for each agent function workflow.add_node("supervisor", team_agents.supervisor_agent) workflow.add_node("researcher", team_agents.researcher_agent) workflow.add_node("writer", team_agents.writer_agent) workflow.add_node("reviewer", team_agents.reviewer_agent) workflow.add_node("reviser", team_agents.reviser_agent) # Set the entry point - always start with the supervisor workflow.set_entry_point("supervisor") # Define conditional edges from the supervisor workflow.add_conditional_edges( "supervisor", lambda state: state["next_action"], # The supervisor's decision dictates the next node { "research": "researcher", "draft": "writer", "review": "reviewer", "revise": "reviser", "finish": END # If supervisor says 'finish', the graph ends } ) # Define regular edges for worker agents to return to the supervisor workflow.add_edge("researcher", "supervisor") workflow.add_edge("writer", "supervisor") workflow.add_edge("reviewer", "supervisor") workflow.add_edge("reviser", "supervisor") # Compile the graph app = workflow.compile() # Optional: Visualize the graph (requires pygraphviz or pydot) # from IPython.display import Image, display # display(Image(app.get_graph().draw_png())) ``` Look at what the conditional edge actually is: a lambda reading one key, and a dict mapping five strings to five node names. That's it. The routing is a lookup, which is why the supervisor's output contract from Step 2 has to hold. **The tell:** if any worker edge points at another worker instead of the supervisor, you've hardcoded a decision the supervisor was supposed to make, and the graph is a chain wearing a costume. --- ### Step 4 — Stream the run so you can watch the state move Compiled graphs stream. Take the stream, because the interesting artifact isn't the final draft, it's the sequence of decisions that produced it. ```python # --- Example Execution --- initial_state: AgentState = { "query": "Explain transformer models in AI.", "research_results": [], "draft": None, "review_comments": [], "iterations": 0, "max_iterations": 2, # Allow up to 2 review/revise cycles "next_action": None, "messages": [] # LangGraph internally uses this for conversational state sometimes } print(f"\n--- Starting Research for: {initial_state['query']} ---") final_state = {} for s in app.stream(initial_state): print(f"Current State: {list(s.keys())[0]} - {s}") final_state.update(s) print("---") print("\n--- Final Output ---") print(f"Topic: {final_state['supervisor']['query']}") print(f"Final Draft:\n{final_state['supervisor']['draft']}") print(f"Total Iterations: {final_state['supervisor']['iterations']}") ``` The output interleaves supervisor decisions with worker prints: delegate to the researcher, back to the supervisor, out to the writer, back, out to the reviewer, and possibly several revise cycles before it finishes. Each agent's print statement shows what it added to the shared state. **The tell:** if the only thing you log is the final draft, you can't distinguish a run that revised twice from a run that skipped review entirely. --- ### Five things to build on top of this Each one is a change to the graph, not to the prompts. - **Real tools in place of the simulated ones.** The researcher currently asks an LLM to recall facts. Swap it for web search, a code interpreter, a database. - **Memory that outlives the run.** The state dict dies when the graph returns. Long-term memory for agents is a different structure. - **A supervisor that spawns workers.** Temporary agents created for niche sub-tasks, rather than a fixed roster of five nodes. - **A workflow that adapts itself.** Routing that changes based on past performance or a new class of problem, instead of five hardcoded rules in a prompt. - **A custom executor in TypeScript.** `type` definitions and interfaces enforce contracts and catch errors before runtime, and a hand-written executor can do real asynchronous parallelism. --- ### What breaks it **The state stops being shared.** If agents start passing context to each other through prompt text rather than through state keys, you're chaining prompts again and calling it a system. The symptom is agents losing context, repeating work, or drifting apart on what the task even is. **The supervisor answers in prose.** `add_conditional_edges` does a dict lookup on `state["next_action"]`. A single word resolves to a node. A polite sentence resolves to nothing. The supervisor's decision quality sets the ceiling for everything downstream, so its output contract is the one prompt in the graph worth over-engineering. **Nothing enforces the iteration cap.** The reviser increments `iterations`, and the supervisor's rules reference `max_iterations`, but that check lives inside a prompt. The graph itself has no counter that stops the review-revise cycle. If the supervisor ignores its own rule four, nothing outside the model is watching. --- ### When a supervisor graph is the wrong choice Most workflows don't need this. Reaching for the machinery when the shape doesn't call for it costs you calls, latency, and new ways to fail. - **The workflow is a straight line.** No branch, no cycle, no conditional. Then it's a sequence of function calls, and the graph is ceremony around three arrows. - **One pass is already good enough.** Every hop is a full LLM call, and the supervisor runs before and after every worker, so it's the most-called node in the graph. You're paying for coordination you didn't need. - **You need actual concurrency.** A synchronous LangGraph setup runs sequentially. The delegation is logical, not parallel, and wall-clock time is the sum of every call in the path. - **You want contracts enforced rather than documented.** `TypedDict` is a hint. Nothing raises when an agent returns a key that isn't in the schema, or omits one that is. That's the reason I'd move the core to TypeScript for anything production-facing. - **You're using the framework as a general orchestration layer.** LangGraph earns its place as a graph-definition DSL for cyclical state machines. Push it past that and you're back to abstraction burying logic you'll want to read. --- ### The shift The framework is the least interesting part of this. What changed how I think about agents is that modeling the problem as a state machine forces explicit answers about roles, responsibilities, and transitions, in a way that a sequential chain lets you avoid. The supervisor is a crude gating network. Each agent is an expert module, activated when the task needs it. That's the MoE idea in miniature, built out of prompts and dict lookups instead of learned routing, and that's why the architecture interests me more than the model behind it. - **Design the state first.** The state dict is the architecture. The agents are functions over it. - **Spend your prompt budget on the supervisor.** It picks every node that runs, so its ceiling is the system's ceiling. - **Make the routing decision a keyword, not a sentence.** An edge is a lookup, not an interpretation. Bigger models aren't the whole story; smarter *teams* of models matter too. > The point is to design intelligence, not just train it. --- Source: https://himanshuat.com/blogs/architecting-multi-agent-systems-with-langgraph ======================================================================== --- title: "Part 4: Architecting Agent Teams - Hierarchical Workflows and Graph Composition" description: "The manager graph needs a conditional edge, and the engine's edge type has no branch for a function. That gap is where a multi-agent system actually gets built - an explicit state contract, a small graph engine, and one node that invokes another graph." date: "March 08, 2024" url: "https://himanshuat.com/blogs/architecting-multi-agent-teams-with-langgraph" --- # Part 4: Architecting Agent Teams - Hierarchical Workflows and Graph Composition The manager graph needs a conditional edge. After a worker finishes it goes back to the router, unless the queue is empty, in which case it synthesizes. The engine's `edges` map is typed `Map`. A function is neither. Neither branch in the execution loop matches, `currentNodeName` never changes, and the same node runs again. > A framework that hides state transitions from you has not removed the complexity. It has removed your ability to see it. That is the whole reason to write the orchestrator yourself. Not because frameworks are bad, but because the routing decision is the system, and you want it in code you can read. ### Chapter 0 — What a graph is here A **node** is an async function from state to state: `(state: WorkflowState) => Promise`. Nothing else. A **graph** is a map of named nodes plus a map of edges and a start node. The engine walks it. The hierarchy comes from one thing only: a node is allowed to call `runWorkflowGraph` on a different graph. There is no separate "team" abstraction, no supervisor class. A manager is a graph whose nodes happen to invoke other graphs. That is the analogy to the brain worth keeping. Cortical regions specialize in vision, language, motor control, or abstract reasoning, and they activate each other through pathways rather than through a central controller. Specialized but integrated. My long-term goal, replicating that through a Mixture of Experts (MoE) architecture, needs exactly this: a router and a set of experts it can call. Flat chains do not get there. A linear workflow is one person handling literature review, data analysis, coding, writing, and editing. It works until the problem has parts, and then it is slow and hard to debug and rarely produces the best result at any stage. --- ### Step 1 — Design the state before you write a node Shared state is the communication bus. Every node reads it, every node writes it, and every graph in the hierarchy sees the same type. Get it wrong and every node downstream inherits the mistake. ```typescript // src/core/workflowState.ts export type AgentStatus = 'PENDING' | 'RUNNING' | 'COMPLETED' | 'FAILED'; export interface WorkflowState { // Global context for the entire workflow overallTask: string; overallResult: string | null; overallStatus: AgentStatus; // Current sub-task being processed by the Manager or a Worker currentSubTask: { id: string; // Unique ID for the sub-task description: string; assignedWorkerGraphId: string | null; // Which worker graph is responsible input: Record | null; // Input payload for the worker output: Record | null; // Output payload from the worker status: AgentStatus; errorMessage: string | null; startTime: number; endTime: number | null; } | null; // History of completed sub-tasks for traceability and error recovery subTaskHistory: Array, 'input'>>; // Exclude large inputs from history // Any global metadata or scratchpad space metadata: Record; // User-specific information or persistent data contextData: Record; } // Initializer for a new workflow state export function createInitialWorkflowState(overallTask: string, contextData: Record = {}): WorkflowState { return { overallTask, overallResult: null, overallStatus: 'PENDING', currentSubTask: null, subTaskHistory: [], metadata: {}, contextData, }; } ``` The load-bearing field is `currentSubTask`. It is a scratchpad with exactly one slot: the manager writes `input`, the worker writes `output`, and the manager reads it back. When it is `null`, nothing is in flight. `subTaskHistory` deliberately drops `input` via the `Omit`, so the audit trail does not carry every payload the workers were handed. **The tell:** if a node reaches for a field the interface does not declare, you do not have a contract yet. You have a bag that nodes agree to be polite about. --- ### Step 2 — Keep the engine small enough to read in one sitting The engine's entire job is: look up a node, run it, look up the next edge, repeat. Everything interesting lives in the nodes. ```typescript // src/core/graphEngine.ts import { WorkflowState } from './workflowState'; // A NodeFunction takes the current state and returns a Promise resolving to the updated state. export type NodeFunction = (state: WorkflowState) => Promise; // A WorkflowGraph defines the nodes and their transitions. export interface WorkflowGraph { id: string; // Unique ID for the graph nodes: Map; // Map of nodeName -> function edges: Map; // Map of nodeName -> nextNodeName(s) or decision logic // Add more complex edge logic if needed (e.g., conditional transitions) startNode: string; } // Global registry for all defined workflow graphs export const graphRegistry = new Map(); // The core engine to execute a workflow graph export async function runWorkflowGraph(graphId: string, initialState: WorkflowState): Promise { const graph = graphRegistry.get(graphId); if (!graph) { throw new Error(`WorkflowGraph with ID '${graphId}' not found.`); } let currentState = { ...initialState }; // Ensure immutability for nodes let currentNodeName = graph.startNode; try { while (currentNodeName) { const nodeFn = graph.nodes.get(currentNodeName); if (!nodeFn) { throw new Error(`Node '${currentNodeName}' not found in graph '${graphId}'.`); } // Mark sub-task as running if applicable if (currentState.currentSubTask && currentState.currentSubTask.assignedWorkerGraphId === graphId && currentState.currentSubTask.status === 'PENDING') { currentState.currentSubTask.status = 'RUNNING'; currentState.currentSubTask.startTime = Date.now(); } console.log(`[${graphId}] Executing node: ${currentNodeName}`); currentState = await nodeFn(currentState); // Execute node, update state const nextEdge = graph.edges.get(currentNodeName); if (!nextEdge) { // End of graph currentNodeName = ''; } else if (typeof nextEdge === 'string') { currentNodeName = nextEdge; } else if (Array.isArray(nextEdge)) { // Example: simple sequential multiple next nodes, or complex conditional routing // For now, let's assume a simple sequential for simplicity or decision-based. // A real implementation would have more sophisticated decision nodes. throw new Error(`Complex edge logic for node '${currentNodeName}' not yet implemented.`); } // A more advanced engine would use a router node to pick from `string[]` or based on `currentState` } // Update overall status if this was the top-level graph if (graphId === currentState.currentSubTask?.assignedWorkerGraphId) { // Check if this was a root worker currentState.currentSubTask.status = 'COMPLETED'; currentState.currentSubTask.endTime = Date.now(); currentState.subTaskHistory.push({ ...currentState.currentSubTask, input: undefined // Clear input before pushing to history to save memory }); currentState.currentSubTask = null; // Clear current sub-task after completion } return currentState; } catch (error: any) { console.error(`[${graphId}] Error during graph execution at node ${currentNodeName}:`, error); if (currentState.currentSubTask && currentState.currentSubTask.assignedWorkerGraphId === graphId) { currentState.currentSubTask.status = 'FAILED'; currentState.currentSubTask.errorMessage = error.message; currentState.currentSubTask.endTime = Date.now(); currentState.subTaskHistory.push({ ...currentState.currentSubTask, input: undefined }); currentState.currentSubTask = null; } currentState.overallStatus = 'FAILED'; currentState.overallResult = `Workflow failed: ${error.message}`; return currentState; } } ``` Two things to notice. The `try/catch` wraps the whole walk, so a throw anywhere inside a node stamps `FAILED` onto the sub-task and returns state rather than propagating an exception to the caller. And `graphRegistry` is a plain module-level `Map`, which means graphs register themselves on import and routing is a string lookup. That string lookup is the seam the hierarchy hangs on. A node that calls `runWorkflowGraph(someId, someState)` is doing the same thing the top-level caller does. **The tell:** if adding a new kind of node means editing the engine, the engine has grown a policy. The engine should not know that managers or workers exist. --- ### Step 3 — Give the manager four jobs and nothing else Decompose, route, invoke, synthesize. The manager never does domain work. It never writes code or analyzes data; it decides who does. ```typescript // src/graphs/mainOrchestratorGraph.ts import { WorkflowState, AgentStatus } from '../core/workflowState'; import { WorkflowGraph, NodeFunction, runWorkflowGraph, graphRegistry } from '../core/graphEngine'; import { generateUniqueId } from '../utils/idGenerator'; // Simple ID generator utility import { llmCall } from '../utils/llmApi'; // Mock LLM API // --- Manager Nodes --- const decomposeTaskNode: NodeFunction = async (state) => { // Use LLM to break down the overall task const prompt = `Given the overall task: "${state.overallTask}", identify the essential sub-tasks required to complete it. Output a JSON array of objects, where each object has 'id', 'description', and 'requiredWorkerGraph' (e.g., 'CodeGenerator', 'DataAnalyzer', 'ReportWriter'). Example: [ { "id": "task_1", "description": "Research market trends", "requiredWorkerGraph": "DataAnalyzer" }, { "id": "task_2", "description": "Generate Python script for analysis", "requiredWorkerGraph": "CodeGenerator" } ]`; const response = await llmCall(prompt, state.contextData); const subTasks = JSON.parse(response); // Assume LLM provides valid JSON // Store decomposed tasks in metadata for sequential processing state.metadata.pendingSubTasks = subTasks; state.overallStatus = 'RUNNING'; return state; }; const expertRouterNode: NodeFunction = async (state) => { if (!state.metadata.pendingSubTasks || state.metadata.pendingSubTasks.length === 0) { // All sub-tasks processed, move to synthesis return { ...state, metadata: { ...state.metadata, nextManagerAction: 'synthesize' } }; } const nextTask = state.metadata.pendingSubTasks.shift(); // Get next task if (!nextTask) throw new Error("No pending sub-tasks, but router node was called."); // Prepare current sub-task for worker invocation state.currentSubTask = { id: nextTask.id, description: nextTask.description, assignedWorkerGraphId: nextTask.requiredWorkerGraph, input: { taskDescription: nextTask.description, context: state.contextData }, // Input for worker output: null, status: 'PENDING', errorMessage: null, startTime: 0, // Will be set by worker graph endTime: null, }; return { ...state, metadata: { ...state.metadata, nextManagerAction: 'invokeWorker' } }; }; const invokeWorkerGraphNode: NodeFunction = async (state) => { if (!state.currentSubTask || !state.currentSubTask.assignedWorkerGraphId) { throw new Error("Attempted to invoke worker without a defined currentSubTask or assignedWorkerGraphId."); } const workerGraphId = state.currentSubTask.assignedWorkerGraphId; console.log(`Manager invoking worker graph: ${workerGraphId} for sub-task: ${state.currentSubTask.id}`); // Create a new state instance for the worker to operate on, containing only relevant input let workerState: WorkflowState = { overallTask: state.currentSubTask.description, // Worker sees its specific task as 'overall' overallResult: null, overallStatus: 'PENDING', currentSubTask: { // This *is* the sub-task for the manager, but the worker treats it as its root id: generateUniqueId(), // Worker gets its own root sub-task ID description: state.currentSubTask.description, assignedWorkerGraphId: workerGraphId, input: state.currentSubTask.input, output: null, status: 'PENDING', errorMessage: null, startTime: Date.now(), endTime: null, }, subTaskHistory: [], // Worker starts with fresh history metadata: {}, contextData: state.contextData, // Pass down relevant context }; // --- CRITICAL: Execute the worker graph as a sub-process --- workerState = await runWorkflowGraph(workerGraphId, workerState); // After worker completes, update manager's state if (!state.currentSubTask) throw new Error("Current sub-task disappeared after worker invocation."); state.currentSubTask.output = workerState.currentSubTask?.output || null; state.currentSubTask.status = workerState.currentSubTask?.status || 'FAILED'; state.currentSubTask.errorMessage = workerState.currentSubTask?.errorMessage || null; state.currentSubTask.endTime = workerState.currentSubTask?.endTime || Date.now(); // Add completed sub-task to manager's history state.subTaskHistory.push({ ...state.currentSubTask, input: undefined // Clear input before pushing to history }); // Clear currentSubTask for the manager to pick the next one state.currentSubTask = null; return state; }; const synthesizeResultNode: NodeFunction = async (state) => { // Collect all outputs from sub-task history const workerOutputs = state.subTaskHistory.map(task => ({ id: task.id, description: task.description, output: task.output, status: task.status })); const prompt = `Synthesize the following sub-task results into a comprehensive final answer for the overall task: "${state.overallTask}" Sub-task results: ${JSON.stringify(workerOutputs, null, 2)} Provide a concise, professional summary.`; state.overallResult = await llmCall(prompt, state.contextData); state.overallStatus = 'COMPLETED'; return state; }; // Define the MainOrchestratorGraph export const MainOrchestratorGraph: WorkflowGraph = { id: 'MainOrchestrator', startNode: 'decomposeTask', nodes: new Map([ ['decomposeTask', decomposeTaskNode], ['expertRouter', expertRouterNode], ['invokeWorkerGraph', invokeWorkerGraphNode], ['synthesizeResult', synthesizeResultNode], ]), edges: new Map([ ['decomposeTask', 'expertRouter'], // Dynamic routing based on 'nextManagerAction' in metadata ['expertRouter', (state: WorkflowState) => state.metadata.nextManagerAction === 'invokeWorker' ? 'invokeWorkerGraph' : 'synthesizeResult'], ['invokeWorkerGraph', 'expertRouter'], // After a worker finishes, go back to router for next task ['synthesizeResult', null], // End of graph ]) as Map string | null)>, // Type assertion for dynamic edges }; graphRegistry.set(MainOrchestratorGraph.id, MainOrchestratorGraph); ``` The topology is a router with a loop hanging off it: ```mermaid flowchart TD A[decomposeTask] --> B{expertRouter} B -->|pending sub-task| C[invokeWorkerGraph] C --> B B -->|queue empty| D[synthesizeResult] ``` That cycle is what a linear chain cannot express, and it is the reason `expertRouterNode` writes `nextManagerAction` into `metadata` instead of returning a next node. The decision has to survive the return so the edge can read it. *A note on dynamic edges:* the `GraphEngine` `edges` definition (`Map`) is basic. For the manager I've shown a more advanced `edges` map with a function that simulates conditional routing, held together by the type assertion at the bottom. A real `GraphEngine` would support this directly through a dedicated `RouterNode` or `ConditionalEdge` type. **The tell:** if a manager node touches worker internals — a code plan, a dataframe, a parsed response — the boundary has already leaked and the worker is no longer swappable. --- ### Step 4 — Hand each worker a fresh state, not a shared one `invokeWorkerGraphNode` does not pass its own state down. It builds a new one. The worker sees a smaller world where its sub-task is the whole job. | field | manager's state | worker's state | |---|---|---| | `overallTask` | the user's original request | the sub-task description | | `currentSubTask.id` | the id from decomposition | a fresh `generateUniqueId()` | | `subTaskHistory` | accumulates every finished sub-task | starts empty | | `metadata` | holds `pendingSubTasks` | starts empty | | `contextData` | from the caller | passed down unchanged | Only `contextData` and the sub-task `input` cross the boundary. That is the isolation: a worker cannot read what other workers did, and it cannot corrupt the manager's queue. ```typescript // src/graphs/codeGenerationGraph.ts import { WorkflowState } from '../core/workflowState'; import { WorkflowGraph, NodeFunction, graphRegistry } from '../core/graphEngine'; import { llmCall } from '../utils/llmApi'; // Mock LLM API import { runPythonCode } from '../utils/codeExecutor'; // Mock code executor // --- Code Generation Worker Nodes --- const planCodeNode: NodeFunction = async (state) => { if (!state.currentSubTask || !state.currentSubTask.input) { throw new Error("Code generation worker requires sub-task input."); } const taskDescription = state.currentSubTask.input.taskDescription; const prompt = `Given the task: "${taskDescription}", outline a Python script plan (steps, libraries, expected output).`; const plan = await llmCall(prompt, state.contextData); state.metadata.codePlan = plan; return state; }; const generateCodeNode: NodeFunction = async (state) => { if (!state.currentSubTask || !state.metadata.codePlan) { throw new Error("Code generation worker needs a plan."); } const taskDescription = state.currentSubTask.input.taskDescription; const plan = state.metadata.codePlan; const prompt = `Based on this plan: "${plan}" and the task: "${taskDescription}", generate the full Python code.`; const code = await llmCall(prompt, state.contextData); state.metadata.generatedCode = code; return state; }; const executeCodeNode: NodeFunction = async (state) => { if (!state.currentSubTask || !state.metadata.generatedCode) { throw new Error("Code execution worker needs generated code."); } const code = state.metadata.generatedCode; try { const executionResult = await runPythonCode(code, state.currentSubTask.input.context); state.currentSubTask.output = { code: code, result: executionResult, success: true, }; } catch (error: any) { state.currentSubTask.output = { code: code, result: null, success: false, error: error.message, }; throw error; // Propagate error for manager to handle } return state; }; export const CodeGenerationGraph: WorkflowGraph = { id: 'CodeGenerator', startNode: 'planCode', nodes: new Map([ ['planCode', planCodeNode], ['generateCode', generateCodeNode], ['executeCode', executeCodeNode], ]), edges: new Map([ ['planCode', 'generateCode'], ['generateCode', 'executeCode'], ['executeCode', null], // End of graph ]), }; graphRegistry.set(CodeGenerationGraph.id, CodeGenerationGraph); // --- Dummy Data Analysis Worker (for illustration) --- const analyzeDataNode: NodeFunction = async (state) => { if (!state.currentSubTask || !state.currentSubTask.input) { throw new Error("Data analysis worker requires sub-task input."); } const taskDescription = state.currentSubTask.input.taskDescription; console.log(`Simulating data analysis for: ${taskDescription}`); // Simulate some work await new Promise(resolve => setTimeout(resolve, 1500)); state.currentSubTask.output = { summary: `Analysis complete for "${taskDescription}". Key finding: Data shows a 15% increase in XYZ.`, rawOutput: { /* large data payload */ }, }; return state; }; export const DataAnalysisGraph: WorkflowGraph = { id: 'DataAnalyzer', startNode: 'analyzeData', nodes: new Map([ ['analyzeData', analyzeDataNode], ]), edges: new Map([ ['analyzeData', null], ]), }; graphRegistry.set(DataAnalysisGraph.id, DataAnalysisGraph); ``` The two workers are different sizes on purpose. `CodeGenerator` plans, generates, then executes, and its `executeCodeNode` both records the failure into `output` and rethrows so the manager sees a `FAILED` status. `DataAnalyzer` is one node. A worker graph is allowed to be trivial; the point is that it is addressable by id. **The tell:** if a worker reads `state.overallTask` expecting the user's original request, it is not a worker. It has been written as a second manager and it will break the moment you call it from somewhere else. --- ### Step 5 — Register everything, then run the top graph Registration happens as a side effect of import. The entry point imports the graph modules purely so their `graphRegistry.set` calls run. ```typescript // src/main.ts import { createInitialWorkflowState } from './core/workflowState'; import { runWorkflowGraph, graphRegistry } from './core/graphEngine'; import { MainOrchestratorGraph } from './graphs/mainOrchestratorGraph'; // This registers itself import { CodeGenerationGraph } from './graphs/codeGenerationGraph'; // This registers itself import { DataAnalysisGraph } from './graphs/dataAnalysisGraph'; // This registers itself async function main() { console.log("Starting hierarchical agent workflow..."); const initialTask = "Analyze sales data for Q3 2023, generate a report, and identify top-performing products. Also, write a Python script to automate weekly sales report generation."; const userContext = { user: "Alice", preferences: { format: "markdown" } }; let state = createInitialWorkflowState(initialTask, userContext); try { state = await runWorkflowGraph(MainOrchestratorGraph.id, state); console.log("\n--- Workflow Completed ---"); console.log("Overall Status:", state.overallStatus); console.log("Overall Result:", state.overallResult); console.log("Sub-task History:"); state.subTaskHistory.forEach(task => { console.log(`- [${task.status}] ${task.description} (${task.assignedWorkerGraphId})`); console.log(` Output: ${JSON.stringify(task.output, null, 2).slice(0, 200)}...`); // Truncate for display }); } catch (error) { console.error("\n--- Workflow Failed ---"); console.error("Final State:", state); console.error("Error:", error); } } main(); // Mock LLM and Code Executor (for demonstration) // In a real system, these would be proper API calls. export const llmCall = async (prompt: string, context: Record): Promise => { console.log(`\n--- LLM Call ---`); console.log(`Prompt: ${prompt.slice(0, 300)}...`); // Simulate LLM processing time await new Promise(resolve => setTimeout(resolve, 500)); if (prompt.includes("identify the essential sub-tasks")) { return JSON.stringify([ { "id": "sub_1", "description": "Analyze Q3 2023 sales data to identify top products", "requiredWorkerGraph": "DataAnalyzer" }, { "id": "sub_2", "description": "Generate Python script for weekly sales report automation", "requiredWorkerGraph": "CodeGenerator" }, { "id": "sub_3", "description": "Write a summary report based on analysis and script", "requiredWorkerGraph": "ReportWriter" } // Assuming another worker ]); } else if (prompt.includes("outline a Python script plan")) { return "Plan: 1. Load sales data. 2. Aggregate by product. 3. Identify top N. 4. Format report. Libraries: pandas, openpyxl."; } else if (prompt.includes("generate the full Python code")) { return "import pandas as pd\\ndef generate_sales_report(data_path):\\n df = pd.read_excel(data_path)\\n # ... (rest of the code)\\n return 'Report Generated'"; } else if (prompt.includes("Synthesize the following sub-task results")) { return `Comprehensive report for Q3 2023:\n1. Top-performing products identified through data analysis.\n2. Automated script for future reports successfully generated and ready for deployment.`; } return "LLM generated a generic response."; }; export const runPythonCode = async (code: string, context: Record): Promise => { console.log(`\n--- Code Execution ---`); console.log(`Executing code: ${code.slice(0, 100)}...`); await new Promise(resolve => setTimeout(resolve, 1000)); return "Code executed successfully. Report data generated."; }; export const generateUniqueId = (): string => `id_${Date.now()}_${Math.random().toFixed(5).replace('0.', '')}`; ``` *(Note: the `ReportWriter` graph is referenced but not fully implemented, to keep the example concise. The `graphRegistry` ensures all workers are known.)* Look at the third element of the mocked decomposition. It routes `sub_3` to `ReportWriter`, and `ReportWriter` is never registered. `runWorkflowGraph` throws `WorkflowGraph with ID 'ReportWriter' not found.` at the moment the manager tries to invoke it, not at startup. **The tell:** if a graph id can appear in a routing decision without existing in `graphRegistry`, your router's output space is larger than your registry, and you find out at runtime. --- ### Four things to build next Everything above is the skeleton. These are the parts I have not built yet, roughly in the order I'd do them. - **LLM-driven routing.** Replace the `if/else` in `expertRouterNode` with a decision that picks worker graphs based on the sub-task's nuances and past performance. That is learning-based routing, straight out of MoE. - **Parallel sub-graphs.** Sub-tasks run strictly sequentially today. Independent ones could run at once, which is both faster and closer to how the brain works. - **Self-healing nodes.** Detect a worker failure and recover: retry the sub-task, escalate to a human, or generate an alternative plan. - **Real tool integration.** Worker graphs need external APIs, databases, and compute, not mocked `llmCall` and `runPythonCode`. --- ### What breaks it **The edge type does not describe the edges.** The manager's `edges` map holds a function and a `null`, and the interface says `string | string[]`. The type assertion at the bottom of `MainOrchestratorGraph` makes it compile, and the engine still has no branch that handles a function. The conditional routing needed care to stay flexible without going opaque, and this is where that care is missing: a `RouterNode` or `ConditionalEdge` type in the engine, not a cast at the call site. **Failure inside a sub-graph does not reach `overallStatus`.** The worker's own `runWorkflowGraph` catches and stamps `FAILED` on its sub-task. `invokeWorkerGraphNode` copies that status into the manager's history and then returns normally, so the router picks up the next task and `synthesizeResultNode` eventually summarizes a set of results that includes failures. Catching a failure deep in a sub-graph and surfacing it back up to the manager is the part you cannot skip, and it does not happen for free just because the state has a `status` field. **State shape leaks.** Early versions of `WorkflowState` were too simplistic and either leaked data across boundaries or left ownership ambiguous. Getting to something that carries both global context and granular sub-task detail took a couple of passes, and every pass touched every node. TypeScript's explicit typing caught most of it at compile time, which is the main argument for typing the state at all. --- ### When a custom graph engine is the wrong choice Most agent work does not need a registry, a manager, or a state machine. Reaching for this when a single pass would do adds ways to fail. - **One specialist, one linear pass.** If there is no branch, no merge, and no cycle, the graph is a function call chain wearing a `Map`. Write the function. - **Sub-tasks that are not independent.** `currentSubTask` is a single slot, filled and cleared one at a time. If task three needs task one's intermediate artifacts and not just its summary, this shape fights you and you will end up stuffing things into `metadata`. - **You need parallelism today.** The engine walks one node at a time and throws on array edges. Fan-out is not a configuration change here; it is a rewrite of the execution loop. - **You need durability across restarts.** State lives in memory, `graphRegistry` is a module-level `Map`, and history discards `input`. Nothing resumes. A framework with checkpointing already solved this and you would be reimplementing it. - **You need the ecosystem more than the control.** Writing your own engine means writing your own tracing, retries, and persistence too. The performance and clarity are real; so is that bill. --- ### The shift The framework complaint that started this — that bloated abstractions dictate too much of *how* you build, hide the underlying logic, and hand you opaque state — only earns its keep if what you write instead is legible. A custom engine you cannot explain is worse than a framework you did not write. What made this legible was not the engine. It was deciding the state contract first and letting everything else be small. - **Design the state before the nodes.** It is the API between every node and every graph, not a data structure. - **Keep the engine ignorant.** It walks nodes and edges. Managers and workers are conventions on top, not concepts inside it. - **Give sub-graphs their own state.** Isolation by construction beats discipline about what nodes are allowed to touch. `MainOrchestratorGraph` is a gating network and each worker graph is an expert; `invokeWorkerGraphNode` is how an expert gets activated. Build the router and the registry first, and the Mixture of Experts is a matter of adding graphs. --- Source: https://himanshuat.com/blogs/architecting-multi-agent-teams-with-langgraph ======================================================================== --- title: "Building Your First RAG System: From Zero to QA Hero" description: "Ask a model about a document it never trained on and it will answer anyway, confidently and wrongly. Retrieval fixes that by handing the model the document at inference time instead of trying to train it in. This post builds the whole pipeline from scratch — loading, chunking, embedding, vector search, and the prompt that is allowed to say it doesn't know." date: "October 14, 2025" url: "https://himanshuat.com/blogs/building-first-rag-system-from-zero-to-qa-hero" --- # Building Your First RAG System: From Zero to QA Hero Ask a model about your company's internal policy from last week. It will answer. It will sound certain. It never saw that document, and nothing in the request told it so. Fine-tuning is the obvious fix and the wrong one. Retraining the model every time a document changes is like rewriting a textbook because a footnote moved. > The model isn't short on knowledge. It's short on a way to look something up. Only one of those is fixed by training. I keep coming back to this because of where it leads. I want to build Mixture of Experts systems where each expert draws on its own tuned memory, and an expert that can't retrieve is just a smaller static snapshot. The retrieval half is the part you can build today, from scratch, in an afternoon. ### Chapter 0 — What RAG actually is Retrieval-Augmented Generation is two pipelines that meet at a store. The first runs once, ahead of time: take a document, cut it into chunks, turn each chunk into a vector, keep the vectors. The second runs per question: turn the question into a vector, find the nearest chunks, paste them into the prompt above the question, and let the model answer from what's in front of it. ```mermaid flowchart LR D[document] --> C[chunks] C --> CE[embed] CE --> S[(vector store)] Q[question] --> QE[embed] QE --> S S --> K[top-k chunks] K --> P[prompt] Q --> P P --> L[LLM] L --> A[answer] ``` The merge at the bottom is the whole idea. The question goes into the prompt twice: once as a vector used to search, once as text the model reads. Everything else is plumbing. Set against the alternative: | | Fine-tuning | Retrieval | |---|---|---| | Adding a document | Retrain | Add vectors to the store | | When facts change | Retrain again | Overwrite the chunk | | Where the fact lives at answer time | In the weights | In the prompt, quotable | | Ongoing cost | Continuous fine-tuning | Embedding plus one call per question | I'll build this in Python. The same six stages port to TypeScript with `node-fetch` and an equivalent embedding library. Raw API calls throughout, no framework wrappers. --- ### Step 1 — Load the document before you complicate it Get the data into a string. That's the entire stage. For this guide it's one text file; in practice you'd pull from a database, an API, PDFs, or Markdown, and the rest of the pipeline wouldn't notice. Here's `my_document.txt`: ```text The quick brown fox jumps over the lazy dog. This is a sample document to demonstrate RAG. RAG systems combine retrieval and generation for better answers. It helps LLMs by providing relevant context. Mixture of Experts (MoE) models can enhance this by routing queries to specialized sub-models. Each expert could have its own RAG system, mimicking specialized cognitive domains. Performance is key in AI systems, demanding efficient data pipelines. ``` ```python # document_loader.py def load_document(file_path: str) -> str: """Loads a text document from the specified path.""" try: with open(file_path, 'r', encoding='utf-8') as f: return f.read() except FileNotFoundError: print(f"Error: Document not found at {file_path}") return "" if __name__ == "__main__": document_content = load_document("my_document.txt") print(f"Loaded document (first 100 chars):\n{document_content[:100]}...") ``` Note what the missing-file path does: it prints and returns an empty string. That's a decision, not an accident. Empty propagates quietly through chunking and embedding, so the guard has to live upstream, which is why the assembled system in Step 6 raises instead of continuing. **The tell:** run the `__main__` block on its own and read the first hundred characters. If you can't see your document there, nothing downstream is worth debugging yet. --- ### Step 2 — Chunk with overlap, then look at the chunks Embedding models and LLMs both have token limits, so a whole book can't be one unit. Cut the text into pieces small enough to embed and large enough to still mean something. The starting strategy is fixed-size chunks with a small overlap, so a sentence sitting on a boundary survives in at least one chunk whole. ```python # chunker.py from typing import List def chunk_text(text: str, chunk_size: int = 200, chunk_overlap: int = 50) -> List[str]: """ Splits text into fixed-size chunks with overlap. A simple, character-based chunker. For production, consider sentence-aware splitting. """ if not text: return [] chunks = [] start = 0 while start < len(text): end = min(start + chunk_size, len(text)) chunk = text[start:end] chunks.append(chunk) if end == len(text): break # Move start position back by overlap for the next chunk start += chunk_size - chunk_overlap # Ensure start doesn't go negative if chunk_size < chunk_overlap if start < 0: start = 0 return chunks if __name__ == "__main__": sample_text = "This is a long sentence that needs to be chunked into smaller pieces. We want to ensure that context is maintained across chunks. Overlap helps with this." chunks = chunk_text(sample_text, chunk_size=50, chunk_overlap=10) print("Generated Chunks:") for i, chunk in enumerate(chunks): print(f"Chunk {i+1}: '{chunk}'") ``` The chunker is *O(N)* in text length, and it's character-based, which means it has no idea what a sentence is. A recursive character splitter, or semantic chunking with spaCy or NLTK, gives better boundaries and costs you complexity. Start with the cheap one. **The tell:** print the chunks and read them. If a fact you'd want to retrieve appears in no chunk in full, your `chunk_size` is wrong for this text, and no amount of embedding quality will recover it. --- ### Step 3 — Embed locally before you pay for embeddings Each chunk becomes a vector. That's what makes similarity searchable: text that means similar things lands in similar places. `all-MiniLM-L6-v2` runs on your machine through `sentence-transformers`, which means iteration costs nothing per run. OpenAI's embedding API is the cloud alternative. Either way, call the library or the API directly rather than through a wrapper like LangChain's `OpenAIEmbeddings`. ```python # embedder.py from typing import List import numpy as np # Prefer a fast, local model for quick iteration. # If you don't have it, run: pip install sentence-transformers from sentence_transformers import SentenceTransformer class Embedder: def __init__(self, model_name: str = 'all-MiniLM-L6-v2'): """ Initializes the embedding model. 'all-MiniLM-L6-v2' is a good balance of speed and performance for many tasks. """ print(f"Loading embedding model: {model_name}...") self.model = SentenceTransformer(model_name) print("Model loaded.") def embed_chunks(self, chunks: List[str]) -> List[np.ndarray]: """ Embeds a list of text chunks into vectors. """ if not chunks: return [] print(f"Embedding {len(chunks)} chunks...") embeddings = self.model.encode(chunks, convert_to_numpy=True) print("Embedding complete.") return embeddings.tolist() # Convert to list of lists/ndarrays for easier storage if __name__ == "__main__": embedder = Embedder() sample_chunks = [ "The quick brown fox jumps over the lazy dog.", "A fast mammal with reddish-brown fur leaps over a sleepy canine." ] embeddings = embedder.embed_chunks(sample_chunks) for i, emb in enumerate(embeddings): print(f"Embedding {i+1} shape: {len(emb)}") print(f"Embedding {i+1} (first 5 values): {emb[:5]}") ``` The two sample sentences share no content words and mean nearly the same thing. That pair is the test. Pass a list rather than looping one string at a time and `SentenceTransformer` batches internally, which is where the throughput on a larger corpus comes from. **The tell:** embed the fox pair and check that they sit closer to each other than to an unrelated line. If they don't, the retrieval stage you're about to build is searching noise. --- ### Step 4 — Store the vectors, then search them honestly Now the vectors need somewhere to live and a way to be searched. A plain Python list with a brute-force scan works and stops working on anything non-trivial. FAISS, Facebook AI Similarity Search, is optimized C++ for nearest-neighbour search and beats a Python loop badly once the dataset grows. `IndexFlatIP` computes inner product. Normalize the vectors with `faiss.normalize_L2` on both sides and inner product becomes cosine similarity, which is the metric you actually want. `IndexFlatL2` is there if you want L2 distance instead. ```python # vector_store.py from typing import List, Tuple import numpy as np # pip install faiss-cpu import faiss class VectorStore: def __init__(self, dimension: int): """ Initializes an in-memory FAISS index. For larger datasets, consider `faiss.IndexFlatL2` for L2 distance, or more advanced indices. """ self.index = faiss.IndexFlatIP(dimension) # IP for Inner Product, suitable for normalized cosine similarity self.texts: List[str] = [] def add_vectors(self, embeddings: List[np.ndarray], texts: List[str]): """ Adds vectors and their corresponding texts to the store. """ if not embeddings or not texts: return if len(embeddings) != len(texts): raise ValueError("Number of embeddings must match number of texts.") embeddings_np = np.array(embeddings).astype('float32') # Normalize embeddings for cosine similarity with Inner Product index faiss.normalize_L2(embeddings_np) self.index.add(embeddings_np) self.texts.extend(texts) print(f"Added {len(embeddings)} vectors to the store.") def search(self, query_embedding: np.ndarray, k: int = 3) -> List[Tuple[str, float]]: """ Searches for the k most similar texts to the query embedding. Returns a list of (text, similarity_score) tuples. """ if self.index.ntotal == 0: return [] query_embedding_np = np.array([query_embedding]).astype('float32') faiss.normalize_L2(query_embedding_np) # Normalize query embedding too distances, indices = self.index.search(query_embedding_np, k) results = [] for i, dist in zip(indices[0], distances[0]): if i != -1: # -1 indicates no result found (shouldn't happen if k <= ntotal) results.append((self.texts[i], dist)) print(f"Retrieved {len(results)} relevant chunks.") return results if __name__ == "__main__": # Example usage (requires an Embedder instance) from embedder import Embedder embedder = Embedder() sample_chunks = [ "The quick brown fox jumps over the lazy dog.", "A fast mammal with reddish-brown fur leaps over a sleepy canine.", "RAG systems combine retrieval and generation for better answers.", "Performance is key in AI systems." ] embeddings = embedder.embed_chunks(sample_chunks) # Initialize VectorStore with the dimension of our embeddings vector_store = VectorStore(dimension=len(embeddings[0])) vector_store.add_vectors(embeddings, sample_chunks) query = "What is RAG?" query_embedding = embedder.embed_chunks([query])[0] search_results = vector_store.search(query_embedding, k=2) print("\nSearch Results for 'What is RAG?':") for text, score in search_results: print(f"Score: {score:.4f}, Text: '{text}'") ``` Read the print line in `search` sceptically. It says "relevant chunks" and it means "the k nearest chunks". Nothing here filters on score. Ask this store about a topic it has never seen and it will return your `k` best matches with a straight face, because that's what nearest-neighbour search does. **The tell:** print the scores, not only the texts. If your top hit for an off-topic question scores about the same as your top hit for an on-topic one, your retrieval isn't discriminating and the prompt is the only thing standing between you and a confident wrong answer. --- ### Step 5 — Write a prompt that is allowed to say no The retrieved context meets the model here. The prompt does three jobs: set the role, hand over the context, ask the question. ```python # llm_client.py import os from typing import List, Dict import openai # pip install openai class LLMClient: def __init__(self, api_key: str, model_name: str = "gpt-3.5-turbo"): """ Initializes the OpenAI LLM client. Ensure OPENAI_API_KEY is set in your environment or passed directly. """ self.client = openai.OpenAI(api_key=api_key) self.model_name = model_name def generate_response(self, prompt_messages: List[Dict[str, str]]) -> str: """ Sends a list of messages to the LLM and returns the generated response. """ try: response = self.client.chat.completions.create( model=self.model_name, messages=prompt_messages, temperature=0.0 # For factual QA, lower temperature is usually better ) return response.choices[0].message.content except openai.AuthenticationError: print("Error: OpenAI API key is invalid or not provided.") return "Error: Could not authenticate with OpenAI. Please check your API key." except Exception as e: print(f"Error calling LLM: {e}") return "Error: Could not generate response." if __name__ == "__main__": # For testing, ensure OPENAI_API_KEY is set in your environment # os.environ["OPENAI_API_KEY"] = "YOUR_API_KEY" api_key = os.getenv("OPENAI_API_KEY") if not api_key: print("Please set the OPENAI_API_KEY environment variable.") else: llm_client = LLMClient(api_key=api_key) test_messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is the capital of France?"} ] response = llm_client.generate_response(test_messages) print(f"LLM Response: {response}") ``` `temperature=0.0` is the one parameter I'd argue about least. This is factual QA against supplied text, and there is nothing to be gained from sampling variety. Calling the API directly is a deliberate choice. Framework layers hide the parameters that matter and add a layer between you and the thing you're trying to profile. If you'd rather run locally, `ollama` or `transformers` slot in behind the same method signature. This is also the only stage that costs money per question, since embedding runs on your own machine. **The tell:** the system message tells the model to say it doesn't have enough information when the context doesn't contain the answer. That instruction is the only refusal mechanism in the whole system. If you never test a question the document can't answer, you don't know whether it works. --- ### Step 6 — Wire it together and keep the seams Six modules, one class. The loader, chunker, embedder, store, and client each stay separately testable, which is the only reason you can swap any one of them later without touching the rest. ```python # rag_system.py import os from document_loader import load_document from chunker import chunk_text from embedder import Embedder from vector_store import VectorStore from llm_client import LLMClient from typing import List, Dict class RAGSystem: def __init__(self, document_path: str, openai_api_key: str): self.document_path = document_path self.embedder = Embedder() self.llm_client = LLMClient(api_key=openai_api_key) self.vector_store: VectorStore = None self._initialize_knowledge_base() def _initialize_knowledge_base(self): """Loads, chunks, and embeds the document to set up the vector store.""" print("Initializing RAG knowledge base...") document_content = load_document(self.document_path) if not document_content: raise ValueError("Failed to load document.") chunks = chunk_text(document_content) if not chunks: raise ValueError("No chunks generated from document.") embeddings = self.embedder.embed_chunks(chunks) if not embeddings: raise ValueError("No embeddings generated.") self.vector_store = VectorStore(dimension=len(embeddings[0])) self.vector_store.add_vectors(embeddings, chunks) print("RAG knowledge base initialized.") def ask(self, query: str, k_retrievals: int = 3) -> str: """ Performs a RAG query: 1. Embeds the user query. 2. Retrieves relevant chunks from the vector store. 3. Constructs a prompt with the retrieved context. 4. Generates a response using the LLM. """ if not self.vector_store: return "RAG system not initialized. Please check document loading." print(f"\nProcessing query: '{query}'") # 1. Embed the user query query_embedding = self.embedder.embed_chunks([query])[0] # 2. Retrieve relevant chunks retrieved_chunks_info = self.vector_store.search(query_embedding, k=k_retrievals) retrieved_texts = [text for text, score in retrieved_chunks_info] # Combine retrieved texts into a single context string context = "\n---\n".join(retrieved_texts) # 3. Construct the prompt system_message = { "role": "system", "content": ( "You are an intelligent QA assistant. " "Use the provided context to answer the user's question. " "If the answer is not in the context, state that you don't have enough information." "Be concise and direct." ) } user_message = { "role": "user", "content": ( f"Context:\n{context}\n\n" f"Question: {query}" ) } prompt_messages: List[Dict[str, str]] = [system_message, user_message] print("Sending prompt to LLM...") # 4. Generate response response = self.llm_client.generate_response(prompt_messages) print("LLM response received.") return response if __name__ == "__main__": # Create a dummy document for demonstration with open("my_document.txt", "w") as f: f.write("""The quick brown fox jumps over the lazy dog. This is a sample document to demonstrate RAG. RAG systems combine retrieval and generation for better answers. It helps LLMs by providing relevant context. Mixture of Experts (MoE) models can enhance this by routing queries to specialized sub-models. Each expert could have its own RAG system, mimicking specialized cognitive domains. Performance is key in AI systems, demanding efficient data pipelines. My favorite fictional character is Sherlock Holmes, a brilliant detective. He uses deductive reasoning to solve complex cases. His methods are a great example of structured problem-solving. """) api_key = os.getenv("OPENAI_API_KEY") if not api_key: print("Please set the OPENAI_API_KEY environment variable to run the RAG system.") else: try: rag_system = RAGSystem(document_path="my_document.txt", openai_api_key=api_key) print("\n--- RAG QA Session ---") queries = [ "What is RAG and how does it help LLMs?", "What is the significance of MoE models in this context?", "Who is Sherlock Holmes?", "What is the capital of Mars?" # This should fail gracefully ] for q in queries: answer = rag_system.ask(q) print(f"\nQuestion: {q}") print(f"Answer: {answer}") except ValueError as e: print(f"Initialization Error: {e}") finally: # Clean up the dummy document if os.path.exists("my_document.txt"): os.remove("my_document.txt") ``` The demo document has Sherlock Holmes in it, which the fox-and-RAG text does not need. He's there so one query has an unambiguous right answer from a completely different part of the document, which tests retrieval rather than the model's general knowledge. The fourth query is the important one. "What is the capital of Mars?" retrieves three chunks, because search always returns `k`, and every one of them is irrelevant. The only thing that stops a confident answer is the sentence in the system message. **The tell:** `_initialize_knowledge_base` runs inside `__init__`. Every restart re-reads, re-chunks, and re-embeds the entire document from zero. If that startup wait annoys you, that annoyance is the argument for a persistent store, arriving on schedule. --- ### Five upgrades to build this week Each of these is a drop-in replacement for one stage, which is the payoff for keeping the modules separate. - **Semantic chunking or a recursive text splitter**, to stop cutting sentences in half at fixed character counts. - **A different embedding model** — open-source ones through `HuggingFaceEmbeddings` for `transformers`, or one you fine-tune yourself. - **A persistent vector store**: `Qdrant`, `Weaviate`, `Pinecone`, or `pgvector` in place of the in-memory index. - **Evaluation metrics** — ROUGE, faithfulness, answer relevancy — so retrieval quality becomes a number instead of an impression. - **A TypeScript port** with `node-fetch` and a server-side embedding service. --- ### What breaks it **Search always returns k.** `IndexFlatIP` has no notion of "nothing here is relevant". Ask about a topic the document doesn't cover and you get your three nearest chunks, scored and formatted exactly like a good retrieval. Every guard against that lives in the prompt, which is a weaker place for it than a score threshold would be. **Cosine similarity depends on normalizing both sides.** The store calls `faiss.normalize_L2` on the corpus in `add_vectors` and again on the query in `search`. Drop either call and the index keeps working, keeps returning results, and quietly stops ranking by the metric you think it's ranking by. Nothing errors. **Fixed-size chunking splits meaning.** A character-based splitter cuts wherever the count lands, mid-sentence, mid-clause. Overlap softens this and doesn't solve it. If the answer to a question straddles a boundary, the retriever may hand the model two halves of a fact and no whole one. **The knowledge base is in memory and rebuilt from scratch.** Kill the process and the index goes with it. Restart and you pay the full embedding cost of the corpus again before answering the first question. --- ### When building RAG from scratch is the wrong choice Retrieval is the right tool for a narrow thing: getting facts the model doesn't have into the prompt. Reach for it outside that and you're carrying five modules for nothing. - **The corpus fits in the prompt.** The demo document is seven lines. Chunking, embedding, and indexing seven lines buys you nothing that pasting them into the system message wouldn't. Build the pipeline when the document is bigger than the window, not before. - **The answer isn't in any document you hold.** Retrieval can only surface what you stored. "What is the capital of Mars?" has no good outcome here; the best case is the model declining to answer. If your users' questions are mostly outside the corpus, you're building an expensive way to say "I don't know". - **You need different behaviour, not different facts.** RAG changes what the model can see. It doesn't change how the model reasons, writes, or formats. If the complaint is about style or method rather than knowledge, retrieval won't touch it. - **You have no way to measure retrieval quality yet.** Nothing in this build tells a good retrieval from a bad one. Until evaluation metrics exist, every tuning decision on chunk size, `k`, or embedding model is a guess dressed as engineering. - **You aren't going to profile it.** The case for raw API calls is control over every parameter and the ability to time each stage. If you won't use that control, a framework's defaults are fine and this is just more code to maintain. --- ### The shift The instinct when a model doesn't know something is to teach it. The cheaper move is almost always to hand it the page. Building this from scratch changes what you can see. Each stage prints what it did, so a bad answer is traceable to a bad chunk, a bad retrieval, or a bad prompt, rather than to a black box with a fluent voice. - **Retrieve before you retrain.** Adding a vector is cheap; adding a fact to the weights is not. - **Trust the scores, not the wording.** Nearest is not the same as relevant, and only one of those is printed by default. - **Keep the seams visible.** Five swappable modules is what makes the upgrade list a week's work instead of a rewrite. This pipeline is also the small version of the thing I actually want: a router sending a query to a domain expert, each expert holding its own retrieval store, specificity and recency at the same time. That starts here, with a document, some chunks, and a prompt that knows how to say no. --- Source: https://himanshuat.com/blogs/building-first-rag-system-from-zero-to-qa-hero ======================================================================== --- title: "Part 2: Building a Self-Correcting RAG with Conditional Edges" description: "A retrieve-then-generate chain has no step whose job is to disagree with the previous one. This is what it takes to add that step — graph state, a critic node, and conditional edges that route the run back to retrieval when the answer is thin." date: "March 15, 2024" url: "https://himanshuat.com/blogs/building-self-correcting-rag-with-langgraph" --- # Part 2: Building a Self-Correcting RAG with Conditional Edges The retriever returns two documents that don't contain the answer. The generator writes a fluent paragraph out of them anyway. The chain returns it and exits. Nothing downstream is in a position to notice, because there is nothing downstream. A linear RAG pipeline retrieves, generates, and stops. > The missing piece in a one-shot RAG isn't a better retriever. It's a step whose only job is to reject the previous step's output and send the run backwards. I want more than a system that answers once. When the answer is weak, it should notice, adjust, and try again. That iterative loop is also how I picture expert modules in an MoE (Mixture of Experts) architecture cooperating and correcting each other. Below is a RAG pipeline built as a self-correcting *graph* rather than a linear chain. I'll use a graph library to show the *concept* of conditional edges, but the mechanics are what matter; in a production system I'd usually reach for raw API calls over a framework. ### Chapter 0 — What a conditional edge is Three primitives, and only the third is new. A **node** is a function that takes the current state and returns an updated copy. An **edge** is a fixed hop: after `retrieve`, always run `generate`. A **conditional edge** is a function of the state that returns the *name* of the next node, evaluated at runtime. That third one is the whole difference. A fixed edge makes a chain. A conditional edge makes a state machine, and a state machine can have cycles — which is another way of saying it can go back and try again. The state is what makes this possible. Routing decisions are functions of the state, so anything the router needs to know has to be *in* the state, not in a local variable in the calling loop. --- ### Step 1 — Put everything the router needs into the state Define what the agent carries through a run before you define any of the steps. A `TypedDict` fits well here: clear and type-safe, with no extra overhead. The fields that matter for routing are `reflection` and `attempts`. The router reads both, so both live in the state alongside the query, the documents, and the answer. ```python from typing import List, Dict, TypedDict, Literal, Optional from langchain_core.messages import BaseMessage, HumanMessage from langchain_openai import ChatOpenAI from langchain_core.documents import Document from langchain_core.output_parsers import StrOutputParser from langchain_core.prompts import ChatPromptTemplate from langgraph.graph import StateGraph, END import os import time # --- Mock/Setup (replace with your actual integrations) --- # For LLM calls, Agno (or a custom httpx client) would be my preference for speed. # Here, using langchain_openai for simplicity in example, but note the intent. os.environ["OPENAI_API_KEY"] = os.getenv("OPENAI_API_KEY", "sk-YOUR-OPENAI-API-KEY") llm = ChatOpenAI(model="gpt-4o-mini", temperature=0) # Fast, cheap for iteration class MockRetriever: """Stands in for a real vector store (Chroma, Qdrant) during development.""" def __init__(self, documents: List[str]): self.documents = documents print("MockRetriever initialized with a simple text store.") def get_relevant_documents(self, query: str) -> List[Document]: # The sleep keeps retrieval latency visible in the per-node timings. time.sleep(0.1) relevant = [ Document(page_content=doc) for doc in self.documents if query.lower() in doc.lower() ] if not relevant: # Never return empty in the demo, or generate() short-circuits to 'fail'. relevant = [Document(page_content=self.documents[i]) for i in range(min(2, len(self.documents)))] print(f"Retrieved {len(relevant)} docs for query: '{query}'") return relevant mock_docs = [ "The quick brown fox jumps over the lazy dog.", "Python is a high-level, interpreted programming language, valued for its readability.", "LangGraph allows building stateful, multi-actor applications with conditional logic.", "Conditional edges in LangGraph enable dynamic routing based on state changes.", "The human brain is an incredibly complex organ, capable of reflection and self-correction.", "Mixture of Experts (MoE) architectures can mimic specialized brain regions, improving scalability.", "Competitive programming emphasizes efficient algorithms and data structures for optimal performance.", "Agno is a lightweight Python client for LLMs, prioritizing raw performance and minimal overhead.", "Neural networks learn through backpropagation, adjusting weights to minimize error.", "RAG pipelines combine retrieval with generation for grounded, accurate answers, reducing hallucinations." ] retriever = MockRetriever(mock_docs) class GraphState(TypedDict): """ Represents the state of our graph. - `query`: The initial or current query. - `documents`: List of retrieved documents. - `answer`: The generated answer. - `reflection`: The result of the reflection step (e.g., "satisfactory", "needs_retrieval", "fail"). - `attempts`: Counter for retry attempts. - `error`: Any error message encountered. """ query: str documents: List[Document] answer: Optional[str] reflection: Optional[Literal["satisfactory", "needs_retrieval", "needs_rewrite", "fail"]] attempts: int error: Optional[str] ``` The state isn't a convenience container. It's the agent's memory for the duration of the run, and the conditional logic falls apart without it — a router that can't see `attempts` cannot decide to stop. **The tell:** if your retry counter lives in the calling loop instead of the state, the routing function is blind to how many attempts already happened, and it will loop forever or not at all. --- ### Step 2 — Give each node exactly one field to write Each node takes the current `GraphState` and returns an updated copy. These nodes are the "experts", specialized functions that each handle one part of the work. `retrieve` writes `documents` and bumps `attempts`. `generate` writes `answer`. `reflect` writes `reflection` and nothing else — it is a pure critic, it does not repair the answer. `rewrite_query` writes `query` and clears `answer` and `documents` so the next cycle starts clean. ```python # Node 1: Retrieval def retrieve(state: GraphState) -> GraphState: """ Retrieves documents based on the current query. """ start_time = time.perf_counter() print(f"---NODE: RETRIEVE (Attempt {state['attempts'] + 1})---") query = state["query"] documents = retriever.get_relevant_documents(query) end_time = time.perf_counter() print(f"Retrieval took {end_time - start_time:.4f} seconds.") return {**state, "documents": documents, "attempts": state["attempts"] + 1} # Node 2: Generation def generate(state: GraphState) -> GraphState: """ Generates an answer using the retrieved documents and the query. """ start_time = time.perf_counter() print("---NODE: GENERATE---") query = state["query"] documents = state["documents"] if not documents: # A deterministic check. No critic call needed to know this run is dead. print("No documents found for generation. Setting reflection to 'fail'.") return {**state, "answer": None, "reflection": "fail"} prompt_template = ChatPromptTemplate.from_messages( [ ("system", "You are a helpful assistant. Provide a concise answer based *only* on the following context. " "If the answer is not in the context, state that clearly and do not make up information.\n\n" "Context:\n{context}"), ("user", "{query}") ] ) rag_chain = prompt_template | llm | StrOutputParser() context = "\n".join([doc.page_content for doc in documents]) answer = rag_chain.invoke({"context": context, "query": query}) end_time = time.perf_counter() print(f"Generation took {end_time - start_time:.4f} seconds.") print(f"Generated Answer: {answer}") return {**state, "answer": answer} # Node 3: Reflection (The "Self-Correction" Brain) def reflect(state: GraphState) -> GraphState: """ Reflects on the generated answer and determines if it's satisfactory or if further action (e.g., re-retrieval, query rewrite) is needed. """ start_time = time.perf_counter() print("---NODE: REFLECT---") query = state["query"] documents = state["documents"] answer = state["answer"] attempts = state["attempts"] if answer is None: print("No answer to reflect upon. Marking as 'fail'.") return {**state, "reflection": "fail"} # The critic's keyword vocabulary is the graph's routing vocabulary. # Change one and you have to change the router. reflection_prompt = ChatPromptTemplate.from_messages( [ ("system", "You are an expert critic. Your task is to evaluate the provided answer based on the original query and context. " "Be strict and objective. Your output should be ONE of the following keywords:\n" "- 'satisfactory': The answer is directly relevant, comprehensive given the context, and addresses the query.\n" "- 'needs_retrieval': The answer is partially relevant but seems to lack sufficient detail or context. More retrieval is needed.\n" "- 'needs_rewrite': The answer is completely off-topic or misunderstands the query. The original query might need reformulation.\n" "- 'fail': The answer explicitly states it cannot find information, or is completely unhelpful/hallucinatory.\n\n" "Original Query: {query}\n" "Retrieved Context:\n{context}\n" "Generated Answer: {answer}\n\n" "Evaluation (single keyword):" ), ("user", "Evaluate the above.") ] ) reflection_chain = reflection_prompt | llm | StrOutputParser() context_str = "\n".join([doc.page_content for doc in documents]) reflection_result = reflection_chain.invoke({ "query": query, "context": context_str, "answer": answer }).strip().lower() # Normalize reflection result to our Literal type if "satisfactory" in reflection_result: reflection_type: Literal["satisfactory", "needs_retrieval", "needs_rewrite", "fail"] = "satisfactory" elif "needs_retrieval" in reflection_result: reflection_type = "needs_retrieval" elif "needs_rewrite" in reflection_result: reflection_type = "needs_rewrite" else: reflection_type = "fail" # Default to fail if LLM gives garbage print(f"Reflection result: {reflection_type}") end_time = time.perf_counter() print(f"Reflection took {end_time - start_time:.4f} seconds.") return {**state, "reflection": reflection_type} # Node 4: Query Rewriting (for iterative improvement) def rewrite_query(state: GraphState) -> GraphState: """ Rewrites the query based on reflection feedback, aiming for better retrieval. """ start_time = time.perf_counter() print("---NODE: REWRITE QUERY---") original_query = state["query"] reflection_result = state["reflection"] answer = state["answer"] if reflection_result == "needs_rewrite": rewrite_prompt = ChatPromptTemplate.from_messages( [ ("system", "The previous retrieval and generation failed to provide a satisfactory answer for the query '{original_query}'. " "The current answer was: '{answer}'. " "Based on this, please reformulate the original query to be more specific or to explore different angles, " "aiming for better retrieval. Provide *only* the new query." ), ("user", "Reformulate the query.") ] ) new_query = (rewrite_prompt | llm | StrOutputParser()).invoke({ "original_query": original_query, "answer": answer }) print(f"Original Query: '{original_query}' -> Rewritten Query: '{new_query}'") return {**state, "query": new_query, "answer": None, "documents": []} # Reset documents and answer for fresh cycle else: # 'needs_retrieval' keeps the query and just re-runs retrieval deeper. print(f"No query rewrite needed. Original query '{original_query}' will be used for re-retrieval.") return {**state, "answer": None, "documents": []} # Reset documents and answer for fresh cycle ``` The LLM appears here as a module, not as the system. It's a critic in one node and a rephraser in another, and the quality of those two prompts sets the ceiling on how well the whole thing self-corrects. That modularity is the same idea underneath MoE. **The tell:** if you can't name the single state field a node exists to write, the node is doing two jobs, and the router will end up depending on a side effect you forgot about. --- ### Step 3 — Route on the state, not on the return value This is where `conditional_edges` earns its keep: the graph picks its own next node from the current state. Re-read the passage, ask a sharper question, or move on — decided at runtime rather than at build time. The topology has one branch and two ways back to the start: ```mermaid flowchart LR R[retrieve] --> G[generate] G --> F{reflect} F -->|satisfactory| E[END] F -->|needs_retrieval| R F -->|needs_rewrite| W[rewrite_query] F -->|fail| E W --> R ``` The critic's four keywords map onto that picture directly: | reflection | what the critic is claiming | next node | |---|---|---| | `satisfactory` | relevant, complete given the context | `END` | | `needs_retrieval` | partially relevant, thin on detail — same query, retrieve again | `retrieve` | | `needs_rewrite` | off-topic or misread the query — the query is the problem | `rewrite_query` | | `fail` | unhelpful, hallucinated, or the critic said something unparseable | `END` | `MAX_ATTEMPTS` sits above all four. The router checks it before it checks the reflection verdict for the retry branches, so a run that keeps returning `needs_retrieval` still terminates. ```python # --- Define the graph --- workflow = StateGraph(GraphState) # Add nodes workflow.add_node("retrieve", retrieve) workflow.add_node("generate", generate) workflow.add_node("reflect", reflect) workflow.add_node("rewrite_query", rewrite_query) # Set entry point workflow.set_entry_point("retrieve") # Define edges workflow.add_edge("retrieve", "generate") workflow.add_edge("generate", "reflect") # The retry budget. Every path back to the start passes through `retrieve`, # which is what increments `attempts`. MAX_ATTEMPTS = 3 def route_reflection(state: GraphState) -> str: print(f"---ROUTING based on Reflection: {state['reflection']}, Attempts: {state['attempts']}---") if state["reflection"] == "satisfactory": return "end" if state["attempts"] >= MAX_ATTEMPTS: print(f"Max attempts ({MAX_ATTEMPTS}) reached. Ending with failure.") return "fail" # Special state to indicate failure or exit point if state["reflection"] == "needs_rewrite": return "rewrite_query" elif state["reflection"] == "needs_retrieval": return "retrieve" # Go back to retrieval with same or potentially deeper query else: # e.g., "fail" or unexpected reflection return "fail" # Add conditional edges from the reflect node workflow.add_conditional_edges( "reflect", route_reflection, { "retrieve": "retrieve", # If needs retrieval, go back to retrieve "rewrite_query": "rewrite_query", # If needs rewrite, go to rewrite_query "end": END, # If satisfactory, end the graph "fail": END # If failed after attempts, end the graph } ) # After rewriting the query, we always go back to retrieval workflow.add_edge("rewrite_query", "retrieve") # Compile the graph app = workflow.compile() # --- Run the graph --- print("\n--- Running the Self-Correcting RAG ---") query_1 = "Tell me about Agno and its purpose." initial_state_1 = GraphState(query=query_1, documents=[], answer=None, reflection=None, attempts=0, error=None) final_state_1 = app.invoke(initial_state_1) print("\n--- Final State 1 ---") print(f"Query: {final_state_1['query']}") print(f"Final Answer: {final_state_1['answer']}") print(f"Reflection: {final_state_1['reflection']}") print(f"Attempts: {final_state_1['attempts']}") print("\n--- Running a challenging query ---") query_2 = "What are the common challenges in MoE architectures that require competitive programming solutions?" initial_state_2 = GraphState(query=query_2, documents=[], answer=None, reflection=None, attempts=0, error=None) final_state_2 = app.invoke(initial_state_2) print("\n--- Final State 2 ---") print(f"Query: {final_state_2['query']}") print(f"Final Answer: {final_state_2['answer']}") print(f"Reflection: {final_state_2['reflection']}") print(f"Attempts: {final_state_2['attempts']}") ``` Note the two runs at the bottom. The first query has a matching document in the corpus. The second deliberately doesn't, which is the case where the loop spends its whole budget and exits on `fail`. Price this before you ship it. Every cycle costs at least two LLM calls, one to generate and one to criticise, plus a third whenever the query gets rewritten, and `MAX_ATTEMPTS = 3` bounds the worst case. A self-correcting answer is a small multiple of a one-shot answer in both latency and spend, not a rounding error on top of it. **The tell:** if you can delete `add_conditional_edges` and the graph still ends the same way on every input you have, there was no branch to begin with — you wrote a chain with extra ceremony. --- ### What breaks it **The critic reads the same context that produced the answer.** `reflect` gets the query, the retrieved documents, and the answer — exactly what `generate` had. A critic with the same evidence as the author tends to agree with the author. The reflection prompt is doing all the work of disagreement, which is why the post's own prompt is written as strictly as it is. **An unparseable verdict is indistinguishable from a real failure.** The normalisation in `reflect` falls through to `fail` when the model returns something outside the four keywords. A garbled critic response and a genuinely hopeless answer produce identical state, and the router treats them identically. If you care about the difference, you need a fifth verdict for "critic malfunctioned" before you can act on it. **Success and exhaustion exit through the same door.** `route_reflection` maps both `"end"` and `"fail"` to `END`. From outside the graph, a good answer and a budget-exhausted answer both look like a completed run. The only way to tell them apart is to read `final_state['reflection']`, which is why the run block prints it. **Latency stacks per node.** Retrieval, generation, criticism, rewriting — each one is a round trip, and the loop multiplies them. Lightweight LLM clients (Agno, or raw `httpx` calls), tighter retrieval queries, and smaller prompts are how you keep the loop inside a budget anyone will tolerate. A loop that takes too long is useless no matter how correct it gets. --- ### When conditional edges are the wrong choice The machinery is only worth it when there's a real decision to make at runtime. Most of the time there isn't. - **Your reject condition is expressible as an `if`.** `generate` already short-circuits on `if not documents` without asking a model anything. If the check you need is "zero results" or "the parse failed", write the branch in Python. A critic call to learn something a boolean already knows is pure cost. - **The corpus doesn't contain the answer.** Rewriting the query and retrieving again re-searches the same documents. That's what `query_2` demonstrates: the loop spends its attempts and exits on `fail`. Retry logic cannot manufacture a fact that was never indexed. - **The pipeline is genuinely linear.** No branch, no merge, no cycle means a `StateGraph`, a `TypedDict`, and a router are ceremony around a function call. In production I'd usually reach for raw API calls over a framework for exactly this reason. - **The latency budget is interactive.** Three cycles at two-to-three model calls each is not something you put behind a search box a user is waiting on. Self-correction belongs where an extra few seconds is cheaper than a wrong answer. --- ### The shift One pass is rarely enough for a hard task, and the fix isn't a better prompt in the generation node. It's a second node that's allowed to say no, plus an edge that points backwards. What that costs you is explicitness. You can no longer keep the run's memory in local variables, because the router has to see it. You can no longer let the loop run until it's happy, because "happy" is a model output and models are happy for bad reasons. - **Put every field the router reads into the state.** A routing decision it can't see is a routing decision it can't make. - **Give the critic its own node and its own vocabulary.** Keywords the router already understands, and nothing else. - **Bound the loop before you build it.** `MAX_ATTEMPTS` first, ambition second. There's a lot left — richer reflection, memory that persists across interactions, MoE setups where expert agents contribute and self-organize. But an agent that can rethink its own answer is already a different class of system from one that answers once and exits. --- Source: https://himanshuat.com/blogs/building-self-correcting-rag-with-langgraph ======================================================================== --- title: "ComfyUI as a research lab" description: "A custom node pack updated, a mask node changed its default on inverted masks, and our outputs moved with nothing changed on our side. Running a try-on diffusion pipeline as a node graph made every experiment a commit, and made very clear where a graph stops being a program." date: "February 04, 2025" url: "https://himanshuat.com/blogs/comfyui-as-a-research-lab" --- # ComfyUI as a research lab A custom node pack updated. A mask node changed its default behaviour on inverted masks. Our outputs shifted, nothing on our side had changed, and the afternoon went into finding out why. Nothing in the workflow JSON records which version of a third-party node interpreted it. That is the shape of every problem I had with ComfyUI. The graph is an excellent record of what you wired. It is silent about everything it depends on. > ComfyUI is a real research environment and a bad engineering environment, and most of the pain comes from expecting one artifact to be both. ### Chapter 0 — What a workflow actually is ComfyUI has a reputation as the thing you use when you don't want to write code. That was not why we used it. A try-on pipeline has two conditioning streams, a segmentation stage, a mask post-process, two merged adapters and a two-pass sampler. The honest question every morning was "which of those did I change yesterday". A node graph makes **topology** the first-class thing. You can't accidentally leave a node connected. The picture on screen is the pipeline, with no gap between what you think is running and what is running. That's the primitive. Not a no-code tool. A serialisable record of what feeds what. --- ### Step 1 — Ask whether you're changing shape or values I started in a notebook, like everyone does. The problem showed up in about two weeks. A diffusion experiment isn't one function you tune. It's a directed graph of components where the interesting variables live at the edges: which encoder feeds which conditioning input, whether the mask is dilated before or after the ROI crop, whether the refinement pass sees the original mask or a shrunk one. In a notebook all of that structure is implicit in the order of cells and the contents of variables that were reassigned six cells ago. When a run comes out good, you have a `.ipynb` where cell 14 was executed after cell 22, and the diff against yesterday is forty lines of JSON with an embedded PNG in it. I could have written a proper config-driven trainer with a YAML per experiment, and for the training side we did. The inference pipeline was different: it was changing shape rather than values, several times a day. | what changes | right artifact | why | |---|---|---| | values, topology fixed | YAML config per experiment | you're sweeping numbers, and a file of numbers diffs cleanly | | shape, several times a day | node graph | rewiring is the experiment, and wiring is what the graph stores | | both, implicitly | notebook | nothing is stored; execution order is the real config | A config file is the wrong artifact when the question is "what if Redux fed the refinement pass but not the primary one". **The tell:** if your last three changes were rewiring rather than retyping a number, the config file is not recording your experiment. --- ### Step 2 — Normalise the export before it reaches git The part that actually changed how we worked is that a ComfyUI workflow serialises to JSON, and JSON goes in git. That sounds trivial and it is not. It means an experiment is a commit. It means "the version that produced the saree renders the client liked" is a SHA rather than a folder called `final_v3_USE_THIS`. It means I can ask what changed between two runs and get an answer in text. The default export fights you. ComfyUI writes node positions, canvas zoom and monotonically increasing node ids into the same file as the semantics, so dragging a node two centimetres to the left produces a diff. We normalised before committing. ```python # normalize_workflow.py """Strip canvas noise out of a ComfyUI workflow export before committing it. Node positions, sizes and the canvas viewport carry no semantics but change constantly, which buries the real diff. This drops them and sorts everything into a stable order so `git diff` shows only what the pipeline does. """ import json import sys from pathlib import Path DROP_KEYS = {"pos", "size", "flags", "order", "extra", "color", "bgcolor"} def normalize(workflow: dict) -> dict: """Return a canvas-independent view of a ComfyUI workflow.""" nodes = [] for node in sorted(workflow.get("nodes", []), key=lambda n: n["id"]): nodes.append( {k: v for k, v in sorted(node.items()) if k not in DROP_KEYS} ) links = sorted(workflow.get("links", []), key=lambda l: (l[1], l[3], l[2])) return { "nodes": nodes, "links": links, "widget_defaults": workflow.get("extra", {}).get("ds", None) and None, } def main(paths: list[str]) -> None: for p in paths: src = Path(p) dst = src.with_suffix(".norm.json") data = json.loads(src.read_text()) dst.write_text(json.dumps(normalize(data), indent=2, sort_keys=True)) print(f"{src.name} -> {dst.name}") if __name__ == "__main__": main(sys.argv[1:]) ``` We ran that as a pre-commit hook, kept the normalised file as the tracked artifact and treated the raw export as a build output. After that, a pull request against the pipeline was readable by a person who hadn't been sitting next to me. **The tell:** if moving a node two centimetres produces a diff, you are committing the canvas, not the pipeline. --- ### Step 3 — Read the graph as topology, then argue about the wiring Here is a stripped-down version of the shape. I want to be clear about what this is: I hand-wrote this fragment for this post. It is not an export from our workflow, our workflow is not published anywhere, and a real ComfyUI file carries a lot more scaffolding than this. Read it as a diagram that happens to be in JSON. ```json { "_file": "flux_vton_plus.illustrative.json", "_note": "hand-written for this post, simplified, not an export", "nodes": [ { "id": 1, "type": "LoadImage", "widgets_values": ["model_0412.png"] }, { "id": 2, "type": "LoadImage", "widgets_values": ["garment_ref_saree_green.png"] }, { "id": 3, "type": "RMBG", "widgets_values": ["u2net", 0.5], "inputs": [{ "name": "image", "link": [2, 0] }] }, { "id": 4, "type": "SAM2Segment", "widgets_values": ["shirt, tshirt, top, saree, dress"], "inputs": [{ "name": "image", "link": [1, 0] }] }, { "id": 5, "type": "GrowMask", "widgets_values": [11, true], "inputs": [{ "name": "mask", "link": [4, 0] }] }, { "id": 6, "type": "StyleModelLoader", "widgets_values": ["flux1-redux-dev.safetensors"] }, { "id": 7, "type": "ReduxApply", "inputs": [ { "name": "style_model", "link": [6, 0] }, { "name": "reference", "link": [3, 0] } ] }, { "id": 8, "type": "UNETLoader", "widgets_values": ["flux1-fill.safetensors", "default"] }, { "id": 9, "type": "LoraLoaderModelOnly", "widgets_values": ["drape_physics_r32.safetensors", 0.6], "inputs": [{ "name": "model", "link": [8, 0] }] }, { "id": 10, "type": "ModelMergeLoRA", "widgets_values": ["occlusion_depth_r32.safetensors", 0.4], "inputs": [{ "name": "model", "link": [9, 0] }] }, { "id": 11, "type": "KSampler", "widgets_values": [30, 3.5, "euler_ancestral", "beta", 1.0], "inputs": [ { "name": "model", "link": [10, 0] }, { "name": "conditioning", "link": [7, 0] }, { "name": "mask", "link": [5, 0] } ] }, { "id": 12, "type": "KSampler", "widgets_values": [10, 3.5, "euler_ancestral", "beta", 0.35], "inputs": [ { "name": "model", "link": [10, 0] }, { "name": "latent", "link": [11, 0] } ] }, { "id": 13, "type": "VAEDecode", "widgets_values": ["fp16"], "inputs": [{ "name": "samples", "link": [12, 0] }] } ] } ``` Three independent branches converge on the first sampler, and the merged model feeds both passes: ```mermaid flowchart LR G[garment ref] --> R[RMBG] R --> RX[Redux apply] SM[style model] --> RX M[model photo] --> S[SAM2 segment] S --> GM[grow mask] U[base UNet] --> L1[drape LoRA] L1 --> L2[occlusion merge] RX --> K1[sampler 30] GM --> K1 L2 --> K1 K1 --> K2[sampler 10] L2 --> K2 K2 --> D[VAE decode] ``` Two things in there are worth pointing at. The RMBG node on the reference garment matters more than it looks. Redux encodes structure from whatever pixels you hand it, so if the reference photo has a mannequin or a studio backdrop in it, some of that leaks into the conditioning and shows up as texture on the output garment. Cutting the garment to a transparent channel first removed a whole category of "why does the sleeve have a shadow that isn't in the reference" bugs. The other is that there are two `KSampler` nodes, at thirty steps and ten steps, and the second runs at low denoise. Fine detail like embroidery and stitching comes out of the second pass. Running the whole thing at forty steps in one pass does not produce the same image and is slower. **The tell:** the RMBG node. If your reference image still has a backdrop in it, some of that backdrop is in your conditioning, and you will spend a day reading it as a model bug. --- ### Step 4 — Put the most-tuned numbers on the canvas The adapters get merged into the base UNet weights rather than swapped at runtime. In the graph that's `model_merge_lora`, and the merge happens once at initialisation, upstream of both samplers. The arithmetic is a linear combination of the low-rank updates: `W_merged = W_base + λ_drape · (B_d A_d) + λ_occ · (B_o A_o)` with λ_drape at 0.6 and λ_occ at 0.4, which is where empirical testing landed. Expert A is draping physics, Expert B is occlusion and depth. These are the most-tuned numbers in the entire pipeline, and they trade against each other in a way you can only see by running it: | pushed too high | what you get | what it costs | |---|---|---| | λ_occ | hands preserved beautifully | garments hang like cardboard | | λ_drape | fabric moves correctly | it moves right over the top of somebody's fingers | The reason this is a graph node rather than a script is that having both coefficients as widget values on a visible node meant anyone could change one, run four images and see the result without touching Python. You can see both numbers and the sampler they feed on one screen. *One caution.* Merging is destructive in the sense that you can no longer attribute a behaviour to one expert once the weights are combined. When output quality regressed, the first debugging step was always to rebuild the graph with a single adapter at a time, which is fast to do in a graph and easy to forget to do. **The tell:** if changing a coefficient requires editing Python, the only person who can run the sweep is the person who wrote the script. --- ### Step 5 — Snapshot everything the graph doesn't name Our convention was one directory per experiment family, normalised JSON tracked, checkpoints referenced by filename and pinned in a lockfile alongside. ```bash # scripts/snapshot_workflow.sh # Freeze the current pipeline: normalised graph + the exact weights it names. set -euo pipefail WORKFLOW="${1:?usage: snapshot_workflow.sh }" TAG="${2:?missing tag}" OUT="experiments/${TAG}" mkdir -p "${OUT}" python scripts/normalize_workflow.py "${WORKFLOW}" cp "${WORKFLOW%.json}.norm.json" "${OUT}/workflow.json" # Record which checkpoint each loader node actually resolved to, by hash. python - "${OUT}/workflow.json" <<'PY' > "${OUT}/weights.lock" import hashlib, json, pathlib, sys MODEL_ROOT = pathlib.Path("/models") graph = json.loads(pathlib.Path(sys.argv[1]).read_text()) for node in graph["nodes"]: for value in node.get("widgets_values", []): if isinstance(value, str) and value.endswith(".safetensors"): path = next(MODEL_ROOT.rglob(value), None) if path is None: print(f"MISSING {value}") continue digest = hashlib.sha256(path.read_bytes()).hexdigest()[:16] print(f"{value} sha256:{digest}") PY # Node pack versions, because a ComfyUI update can silently change behaviour. pip freeze | grep -Ei "torch|diffusers|transformers" > "${OUT}/python.lock" (cd custom_nodes && for d in */; do printf "%s %s\n" "${d%/}" "$(git -C "$d" rev-parse --short HEAD 2>/dev/null || echo unpinned)" done) > "${OUT}/nodes.lock" git add "${OUT}" && git commit -m "snapshot: ${TAG}" echo "snapshotted ${TAG}" ``` That `nodes.lock` file exists because of the afternoon this post opens with. Pinning the node repos by commit is the only defence, and I'd do it on day one next time rather than after losing a day. **The tell:** if the outputs moved and `git diff` is empty, the thing that changed was never in the file. --- ### Four things to set up on day one - **A defined moment where a workflow graduates.** We eventually reimplemented the settled parts of the pipeline as a Python service and kept ComfyUI for exploration. Doing that split deliberately and early would have saved a rewrite. - **Pinned third-party node packs, from the first commit.** The failure mode where nothing you changed changed the output is expensive to diagnose and trivial to prevent. - **A one-paragraph README next to every experiment directory.** The graph records what ran. It records nothing about why, and three weeks later "why" is the only thing you actually want. - **The expectation that the artifact is a diagram, not a program.** Once I started treating the workflow JSON as an executable diagram with a lockfile stapled to it, my expectations matched reality and I stopped being annoyed at it. --- ### What breaks it **A dependency changes underneath the file.** The workflow JSON pins the topology and the widget values. It does not pin the node implementations, the ComfyUI version, the model file contents or the CUDA and PyTorch stack underneath. Everything above about lockfiles is scaffolding bolted on around a format that was not designed to carry that information. **The merge erases attribution.** Two adapters combined into the base weights is one set of weights. Nothing in the graph tells you which expert produced the behaviour you're looking at, and the rebuild-with-one-adapter step is the easiest thing in this whole process to skip when you're in a hurry. **Duplication drifts.** You cannot easily write the graph equivalent of a helper function that you call from four places. Subgraphs help and are not the same thing. Our workflow files had duplicated sections that drifted apart, which is exactly the failure mode you'd predict from a language without abstraction. --- ### When the graph is the wrong choice I'll defend "ComfyUI is a real research environment". I won't defend "ComfyUI is a good engineering environment". These are the cases where the second claim is the one that matters. - **The topology is fixed and you're sweeping values.** A config file is the right artifact here, and we used one for the training side. The graph buys you nothing when nothing is being rewired. - **The change has to survive asynchronous review.** A reviewer looking at `"widgets_values": [11, true]` changing to `"widgets_values": [7, true]` has no idea that this is the mask dilation radius unless they already know the pipeline. No function names, no types, no comments and nowhere to put one. Every meaningful review we did happened by standing at a screen together, which does not scale past the people in the room and leaves no record. - **The output reaches a customer.** Repeatability is partial by construction, so anything on the serving path should be exported to code first. A normalised JSON plus three lockfiles is a good approximation of reproducibility and not the real thing. - **You need assertions on intermediate state.** There is no check you can attach to a node saying the mask must cover between four and sixty percent of the frame. That check lives in Python in the serving path, which means the research pipeline and the production pipeline validate different things. --- ### The shift The graph earned its place by making topology reviewable in a git diff. It never earned the place I kept trying to give it, which was source code. Every hour I lost went into that gap: a node pack that changed, a widget value nobody could interpret, a subgraph that drifted from its copy. - **Store shape, not values.** If you're rewiring daily, the wiring is the experiment and it belongs in the tracked artifact. - **Pin what the file doesn't name.** Node pack commits, weight hashes, the Python stack. The workflow records none of it. - **Graduate the settled parts to code.** Exploration in the graph, serving in Python, with an explicit moment where one becomes the other. The next thing I want to build is the bridge: a small exporter that walks a normalised workflow and emits the equivalent diffusers calls, so graduating an experiment is a mechanical step rather than a careful manual transcription that quietly drops a mask dilation somewhere. --- Source: https://himanshuat.com/blogs/comfyui-as-a-research-lab ======================================================================== --- title: "Composing adapters that disagree with each other" description: "Two LoRA experts merged with two scalars shipped fine. A whole library of them for demographics, poses, lighting and backgrounds did not, because adapters conflict per layer and one blend weight per adapter cannot express that." date: "August 05, 2025" url: "https://himanshuat.com/blogs/conflict-aware-adapter-composition" --- # Composing adapters that disagree with each other Two adapters, two scalars, one merged checkpoint. That shipped. The same trick on a library of adapters gave me a model that was good almost everywhere and wrong in a few specific places, and the one knob I had could only make it worse everywhere in exchange for fixing those places. > A global blend weight is a per adapter instrument applied to a per layer disagreement. This post is the shape of that problem and what we ended up serving. It stops short of the merging method itself, and I say why partway down. ### Chapter 0 — What a merge is when it works The try-on model fused two experts into the base weights with one coefficient each: $$ W_{\text{merged}} = W_{\text{base}} + \lambda_{\text{drape}}(B_d A_d) + \lambda_{\text{occ}}(B_o A_o) $$ We settled on `lambda_drape = 0.6` and `lambda_occ = 0.4` by testing, which is a polite way of saying we swept a two dimensional grid and looked at pictures until one cell stopped being objectionable. That works when the two adapters were trained on problems that barely touch each other. Fabric draping and hand occlusion overlap a little, not a lot, and 0.6 against 0.4 was enough to keep both. The base image stage is not that. Before any garment gets composited onto anything, we generate the scene: a model of a particular demographic, in a particular pose, under a particular lighting setup, against a particular backdrop. Each of those is an adapter, some trained in house on our curated corpus and some pulled from good public work. A brand will ask for a combination nobody ever trained together. The coefficient grid is now too large to sweep by hand. Sweeping it harder wouldn't help, because the coefficient is the wrong instrument. --- ### Step 1 — Find out where two adapters even touch A LoRA is a set of low rank updates attached to specific modules inside the transformer. Two adapters trained separately land on overlapping sets of modules. Where they overlap, their updates may point in compatible directions or in opposing ones, and that varies module by module inside the same pair of adapters. Here's what that means in practice. A demographics adapter carries a great deal of information about skin: undertone, pore structure, how highlights sit on a cheekbone. A high key lighting adapter also has opinions about skin, because that's most of what high key lighting is doing to a portrait. Those two agree about the backdrop and fight about the face. Turn the lighting adapter down to protect skin tone and you lose the backdrop falloff it was rendering perfectly well. Turn it up and the face goes waxy. So the first thing I built was not a merge. It was a diagnostic that reports where two adapters land on the same modules. ```python # tools/inspect_adapter_overlap.py """Report which modules two LoRA adapters both write to. Overlap on its own is not evidence of conflict. It tells you where a conflict is possible, which is the only thing a static check of two files on disk can honestly tell you. Direction is a separate question and needs the model in memory. """ from __future__ import annotations import sys from collections import OrderedDict from safetensors.torch import load_file def target_modules(path: str) -> "OrderedDict[str, int]": """Map module name to the rank of its down projection.""" tensors = load_file(path) modules: "OrderedDict[str, int]" = OrderedDict() for key, tensor in tensors.items(): if not key.endswith("lora_down.weight"): continue name = key.rsplit(".lora_down", 1)[0] modules[name] = tensor.shape[0] return modules def overlap_report(path_a: str, path_b: str, preview: int = 12) -> None: """Print the shared and exclusive module sets for two adapters.""" a = target_modules(path_a) b = target_modules(path_b) shared = [m for m in a if m in b] only_a = [m for m in a if m not in b] only_b = [m for m in b if m not in a] print(f"A: {path_a} modules={len(a)}") print(f"B: {path_b} modules={len(b)}") print(f"shared={len(shared)} a_only={len(only_a)} b_only={len(only_b)}") print() for name in shared[:preview]: print(f" shared {name} rank_a={a[name]} rank_b={b[name]}") for name in only_a[:preview]: print(f" a_only {name} rank={a[name]}") if __name__ == "__main__": overlap_report(sys.argv[1], sys.argv[2]) ``` Running that across the adapter library is unglamorous and it changed how I thought about the merge. The overlap is never total and never trivial. Adapters that feel semantically unrelated share a surprising number of attention modules, and adapters that feel like near duplicates diverge in the blocks that carry composition. Once you have that laid out per module, a single coefficient stops looking like a simplification and starts looking like a category error. **The tell:** if two adapters you'd have called unrelated turn out to share attention modules, your blend weight is arbitrating a fight you never knew was happening. --- ### Step 2 — Write the priors down, in a file, before you write the merge The thing that decides which adapter should win where is not a metric. It's a decision about what the image is for. On an e-commerce shoot, skin follows the demographics adapter and nothing else gets to touch it, because that's the part a brand's legal and marketing teams will both look at. The backdrop follows the background adapter. Garment surface rendering follows lighting. Those are priorities somebody with taste sets, writes down, and defends in a review. So they live in a file, next to the model they produce. ```yaml # priors/editorial_ecom.yaml # Human authored, reviewed by whoever owns the brand look, versioned with # the checkpoint it produces. The resolver consumes this file. It never # writes it, and it is not allowed to reorder the priority block. base: flux adapters: - id: demographics_in_female_25_35 origin: in_house rank: 32 - id: pose_studio_standing origin: in_house rank: 32 - id: lighting_softbox_high_key origin: public rank: 32 - id: background_studio_sweep origin: public rank: 32 # Which adapter owns which visual concern when two of them disagree. priority: skin_and_hair: demographics_in_female_25_35 body_proportion: pose_studio_standing garment_surface: lighting_softbox_high_key frame_and_backdrop: background_studio_sweep # Concerns that no other adapter may override, regardless of what the # resolver would otherwise prefer. protected: - concern: skin_and_hair from: [lighting_softbox_high_key, background_studio_sweep] output: name: aurax-v1 bake: true ``` This is the artifact I was most wrong about going in. I expected the valuable output of the project to be an algorithm. The algorithm matters, and the priors file is what makes the algorithm produce something a brand signs off on. Two brands with identical adapter libraries and different priors files get visibly different models. **The tell:** when a shoot came back with notes, nine times out of ten the fix was in the priors and not in the code. If your fixes keep landing in code, you haven't written the priors down yet. --- ### Step 3 — Bake one checkpoint instead of stacking at request time The other decision worth writing down is that the output of composition is a single checkpoint, which we called AuraX-V1, and not a set of adapters loaded at request time. Runtime stacking is the obvious design and it's tempting because it keeps everything composable. It also means every request carries the low rank matmuls for every active adapter in the hot path, holds each adapter resident, and exposes a serving surface where any caller can request any combination, including combinations nobody has ever looked at. | | runtime stacking | build time bake | |---|---|---| | where composition happens | request path | build path | | what serving loads | every active adapter, resident | one file | | what a cold worker pulls | a graph to assemble | one checkpoint | | what you can review | the assembling code | the model itself | | roll back unit | none | one version tag | | per request composability | yes | no | That fourth row is the real cost. If the set of possible models is generated at request time, you cannot review the model, you can only review the code that assembles it. ```mermaid flowchart LR A[adapters] --> R[resolve] P[priors file] --> R B[base weights] --> R R --> C[one checkpoint] C --> S[serving] A -.-> K[runtime stack] K -.-> D[R&D only] ``` The solid path is production. The dotted one is the stacking path we kept for research. ```python # pipelines/bake_runtime_model.py """Resolve an adapter set against a priors file and bake one checkpoint. The output of this script is the artifact that serving loads: one file, one version tag, one thing to roll back. Composition happens here, at build time, never in the request path. """ from __future__ import annotations import argparse import json from pathlib import Path import torch import yaml from safetensors.torch import load_file, save_file # Internal. This is the part of the system the blog post does not open up. from aurax.caac import resolve # noqa: F401 def load_priors(path: Path) -> dict: """Read and lightly validate a priors file.""" priors = yaml.safe_load(path.read_text()) declared = {a["id"] for a in priors["adapters"]} referenced = set(priors["priority"].values()) missing = referenced - declared if missing: raise ValueError(f"priority references unknown adapters: {sorted(missing)}") return priors def bake(priors_path: Path, base_path: Path, out_dir: Path) -> Path: """Produce the fused runtime checkpoint and its provenance sidecar.""" priors = load_priors(priors_path) base = load_file(str(base_path)) adapters = { a["id"]: load_file(f"adapters/{a['id']}.safetensors") for a in priors["adapters"] } merged = resolve(base=base, adapters=adapters, priors=priors) out_dir.mkdir(parents=True, exist_ok=True) ckpt = out_dir / f"{priors['output']['name']}.safetensors" save_file(merged, str(ckpt)) sidecar = { "name": priors["output"]["name"], "base": priors["base"], "adapters": [a["id"] for a in priors["adapters"]], "priors_sha": priors_path.read_bytes().hex()[:16], "torch": torch.__version__, } (out_dir / "provenance.json").write_text(json.dumps(sidecar, indent=2)) print(f"wrote {ckpt}") print(f"fused {len(adapters)} adapters into one checkpoint") return ckpt if __name__ == "__main__": ap = argparse.ArgumentParser() ap.add_argument("--priors", type=Path, required=True) ap.add_argument("--base", type=Path, required=True) ap.add_argument("--out", type=Path, default=Path("build/")) args = ap.parse_args() bake(args.priors, args.base, args.out) ``` The cost is real. You lose per request composability, and every new priors file means a new bake and a new review. We took that trade because brands don't change their house look between requests. For R&D we kept the stacking path, since being able to hot swap an adapter is worth a lot when you're still deciding what the adapter should do. **The tell:** ask what you would roll back to if tomorrow's images came out wrong. If the answer is a version tag on one file, you baked. If it's a diff in an assembler, you didn't. --- ### The part I'm not opening up The `resolve` call above is the method, and I'm not publishing it. That covers what it does internally and the index we use to talk about how much two adapters interfere. It's company IP and it's the reason a brand pays us instead of running `model_merge_lora` themselves. I'd rather say that plainly than write a paragraph that gestures at a method without being reproducible. Vague method sketches cost the reader time, can't be checked, and imply a rigour the writing isn't carrying. The problem statement and the outcome are the parts that are useful to anyone outside the company anyway, and both are here. --- ### Step 4 — Decide up front whether the comparison is qualitative We compared the fused model against base Flux-dev, Google's Imagen 4, ChatGPT (August) and Nano-banana, on skin texture and commercial realism, by looking at the images. No score, no leaderboard. That was a deliberate choice and it's also a weakness, so both halves are worth stating. We didn't score it because the generic aesthetic scorers available to us were not measuring what we needed. They lean toward darker, moodier imagery, which is a defensible notion of aesthetic and a bad fit for an e-commerce product shot that has to be evenly lit and honest about the garment. We ended up building our own human-in-the-loop scorer for exactly this reason, but for this comparison we used our eyes. | model | what the images showed | |---|---| | fused (AuraX-V1) | natural skin and hair, clean well lit compositions | | Flux-dev | dramatic contrast, good for a campaign, competes with the garment | | Imagen 4 | photorealism reads slightly too perfect on close inspection | | Nano-banana | strong contrast some brands will love and others reject on sight | That is a qualitative judgment made by people with a commercial stake in the answer, and you should read it as such. All of it came out of a fine-tune on a comparatively modest corpus of **5,000 curated images**, which is the number that surprised me most. **The tell:** if your scorer prefers darker and moodier and your product has to be evenly lit, the scorer is not measuring your product. Say the comparison is qualitative and run it properly, or don't run it. --- ### Four things to get right on the next library - **Write the priors before training the adapters.** We did it in the other order and spent weeks training capability we then had to suppress. - **Build the overlap diagnostic on day one.** It's forty lines, it needs no GPU, and it reframes the problem faster than any amount of staring at outputs. - **Keep the runtime stacking path alive for research** even after you commit to baking for production. We nearly deleted it and it's where every subsequent experiment started. - **Commit to a blind panel if the comparison is going to be qualitative.** Ours drifted into an informal one, and that limits how hard I can lean on it. --- ### What breaks it **Training before the priors exist.** Adapter capability you didn't ask for is capability the resolver has to suppress later. We paid weeks for that ordering mistake, and the priors file was the cheap artifact the whole time. **Judging with a scorer that has different taste than your product.** The generic aesthetic scorers we could reach reward darker and moodier images. Point one at an e-commerce shot that has to be evenly lit and it will confidently grade the wrong axis. That's what pushed us to build a human-in-the-loop scorer instead. **An informal comparison presented as a result.** Ours was people with a commercial stake looking at pictures. It was enough to decide what to ship and it is not enough to claim a win over Imagen 4. Those are different sentences and the informal method only supports one of them. --- ### When per layer composition is the wrong choice Most merges do not need any of this. The machinery costs a review cycle per priors file and a person with the authority to write one. - **Your adapters barely overlap.** Fabric draping and hand occlusion touch a little. Two scalars at 0.6 and 0.4 held both, and we shipped that. If a grid sweep finds a cell you stop objecting to, take it. - **The house look changes between requests.** Baking assumes the set of models is small and reviewable. If callers legitimately need arbitrary combinations at request time, you want the stacking path and its costs, not a checkpoint per combination. - **You're still deciding what the adapter should do.** Composition presumes the pieces are stable. During R&D, hot swapping an adapter is worth more than a reviewable artifact, which is why we kept that path. - **Nobody owns the priorities.** The resolver's input is a human decision about which adapter owns skin. Without someone who will write that down and defend it in a review, you've automated the arbitration of a question nobody has answered. --- ### The shift I went in expecting the valuable output to be an algorithm. What made brands sign off was a YAML file a human wrote, and the algorithm was the thing that executed it faithfully. The other correction: the first useful artifact was a forty line diagnostic with no GPU in sight. I built it late and it changed the problem statement. - **Look at module overlap before you pick a coefficient.** The conflict is per layer whether or not your instrument is. - **Put the priorities in a versioned file, not in the merge code.** Nine fixes out of ten land there. - **Ship one checkpoint you can review and revert.** A model assembled per request is a model nobody has seen. Next on my list is measurement, specifically the try-on side, where we do have numbers and where I've started to suspect the aggregate ones are telling us less than they appear to. --- Source: https://himanshuat.com/blogs/conflict-aware-adapter-composition ======================================================================== --- title: "Beyond Pre-builts: Crafting Custom Tools for Domain-Specific LangChain Agents" description: "A generic pre-built tool makes the agent guess the function, the arguments, and their shape, and guessing is where hallucinations come from. Custom tools replace the guess with a typed contract the model can read." date: "March 22, 2024" url: "https://himanshuat.com/blogs/crafting-custom-tools-for-langchain-agents" --- # Beyond Pre-builts: Crafting Custom Tools for Domain-Specific LangChain Agents Point an agent at an inventory system with nothing but a generic `SearchTool` and a `Calculator`, and ask it how many laptops are in warehouse W01. It has to guess. What the operation is called, what arguments it takes, what shape they arrive in. Parsing complex language into an unstructured call is exactly where you get hallucinations and poor performance out of a model that could have done the job. > The agent is an orchestrator, not a worker. It turns natural language into a structured function call. Everything that determines whether the system works lives in the tools you write. ### Chapter 0 — What a tool actually is A tool is any function, class method, or API call, presented to the LLM as an unambiguous, type-safe interface. That is the whole primitive. The hard part is not writing the function, it's presenting it. What draws me to AI isn't building the next big model, it's understanding and replicating biological intelligence, and the brain is not a monolithic LLM. It's a specialized Mixture-of-Experts architecture: different cortical areas handle vision, language, motor control, and memory, each tuned to its domain. A raw LLM is capable but generalist. Tools are the specialized organs you give it. So the tools you want are not `SearchTool` and `Calculator`. They're bespoke: * `query_inventory(product_id: str)` * `fetch_stock_price(ticker: str, date: str)` * `activate_robot_arm(position: List[float])` LangChain gives you two mechanisms to expose these. | | `Tool` | `StructuredTool` | |---|---|---| | Input | usually a single string | multiple or complex arguments | | Validation | whatever you write yourself | Pydantic, via `args_schema` | | Reach for it when | the call really is one string | anything serious | --- ### Step 1 — Define the input schema before you write the logic Write the contract first. Type hints and docstrings become a JSON schema the model's function-calling consumes directly, which is what removes the guessing. ```python # tools/schemas.py from pydantic import BaseModel, Field from typing import List, Optional class StudentGradeInput(BaseModel): """Input for retrieving a student's grade.""" student_id: str = Field(description="The unique identifier for the student.") course_name: Optional[str] = Field(None, description="The name of the course to get the grade for. If not provided, returns all grades.") class UpdateStudentGradeInput(BaseModel): """Input for updating a student's grade in a specific course.""" student_id: str = Field(description="The unique identifier for the student.") course_name: str = Field(description="The name of the course to update the grade for.") new_grade: float = Field(description="The new numerical grade for the student in this course.") class InventoryQueryInput(BaseModel): """Input for querying product inventory information.""" product_id: str = Field(description="The unique identifier for the product.") warehouse_id: Optional[str] = Field(None, description="The specific warehouse to check inventory in. Defaults to all warehouses if not provided.") class AddInventoryInput(BaseModel): """Input for adding new inventory to a product.""" product_id: str = Field(description="The unique identifier for the product.") quantity: int = Field(description="The quantity to add to the product's inventory.") warehouse_id: str = Field(description="The specific warehouse where inventory is being added.") ``` Every `Field` description is written for the model, not for your teammates. `course_name` being optional is not a detail the model can infer from the function body it never sees. **The tell:** hand the schema alone to someone who has never seen your codebase and ask them to construct a valid call. If they have to ask you a question, the model is guessing too. --- ### Step 2 — Write the backend as if no agent existed The logic is a normal function. No LangChain import, no agent-shaped concessions, nothing that assumes an LLM is on the other end. Here that means a local SQLite database and a private HTTP API. Note the directness: no ORMs, just raw `sqlite3` where performance matters, and `requests` for HTTP. ```python # tools/backend.py import sqlite3 import requests import json import os from typing import Dict, Any, List, Optional # --- Database Mock-up --- DB_NAME = "school_db.sqlite" def _init_db(): conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS students ( id TEXT PRIMARY KEY, name TEXT ) """) cursor.execute(""" CREATE TABLE IF NOT EXISTS grades ( student_id TEXT, course_name TEXT, grade REAL, PRIMARY KEY (student_id, course_name), FOREIGN KEY (student_id) REFERENCES students(id) ) """) # Seed data cursor.execute("INSERT OR IGNORE INTO students (id, name) VALUES (?, ?)", ('S001', 'Alice')) cursor.execute("INSERT OR IGNORE INTO students (id, name) VALUES (?, ?)", ('S002', 'Bob')) cursor.execute("INSERT OR IGNORE INTO grades (student_id, course_name, grade) VALUES (?, ?, ?)", ('S001', 'Math', 95.5)) cursor.execute("INSERT OR IGNORE INTO grades (student_id, course_name, grade) VALUES (?, ?, ?)", ('S001', 'Physics', 88.0)) cursor.execute("INSERT OR IGNORE INTO grades (student_id, course_name, grade) VALUES (?, ?, ?)", ('S002', 'Math', 72.0)) conn.commit() conn.close() _init_db() # Ensure DB is initialized on import def get_student_grade(student_id: str, course_name: Optional[str] = None) -> Dict[str, Any]: """Retrieves a student's grade(s) from the local database.""" conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() if course_name: cursor.execute("SELECT grade FROM grades WHERE student_id = ? AND course_name = ?", (student_id, course_name)) result = cursor.fetchone() conn.close() return {"student_id": student_id, "course_name": course_name, "grade": result[0] if result else None} else: cursor.execute("SELECT course_name, grade FROM grades WHERE student_id = ?", (student_id,)) results = cursor.fetchall() conn.close() return {"student_id": student_id, "grades": [{ "course_name": r[0], "grade": r[1] } for r in results]} def update_student_grade(student_id: str, course_name: str, new_grade: float) -> Dict[str, Any]: """Updates a student's grade in a specific course in the local database.""" conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute(""" INSERT OR REPLACE INTO grades (student_id, course_name, grade) VALUES (?, ?, ?) """, (student_id, course_name, new_grade)) conn.commit() conn.close() return {"status": "success", "student_id": student_id, "course_name": course_name, "new_grade": new_grade} # --- Private API Mock-up --- # In a real scenario, this would be an actual API endpoint. # For demonstration, we'll simulate a simple in-memory "API" _INVENTORY_DATA = { "P001": {"name": "Laptop", "stock": {"W01": 15, "W02": 10}}, "P002": {"name": "Mouse", "stock": {"W01": 50, "W03": 25}}, } API_BASE_URL = "http://mock-inventory-api.com/api/v1" # This URL is purely illustrative def _mock_api_call(method: str, path: str, json_data: Optional[Dict[str, Any]] = None) -> Dict[str, Any]: """Simulates an API call to our private inventory system.""" print(f"DEBUG: Mock API call - {method} {path} with data: {json_data}") # In a real app, you'd use requests.request(...) if path == "/inventory/query": product_id = json_data.get("product_id") warehouse_id = json_data.get("warehouse_id") if product_id not in _INVENTORY_DATA: return {"error": "Product not found", "status_code": 404} product_info = _INVENTORY_DATA[product_id] if warehouse_id: if warehouse_id not in product_info["stock"]: return {"error": f"Warehouse {warehouse_id} not found for product {product_id}", "status_code": 404} return {"product_id": product_id, "name": product_info["name"], "warehouse_id": warehouse_id, "stock": product_info["stock"][warehouse_id]} else: return {"product_id": product_id, "name": product_info["name"], "stock_by_warehouse": product_info["stock"]} elif path == "/inventory/add" and method == "POST": product_id = json_data.get("product_id") quantity = json_data.get("quantity") warehouse_id = json_data.get("warehouse_id") if product_id not in _INVENTORY_DATA: _INVENTORY_DATA[product_id] = {"name": f"New Product {product_id}", "stock": {}} # Auto-create product for demo if warehouse_id not in _INVENTORY_DATA[product_id]["stock"]: _INVENTORY_DATA[product_id]["stock"][warehouse_id] = 0 _INVENTORY_DATA[product_id]["stock"][warehouse_id] += quantity return {"status": "success", "product_id": product_id, "quantity_added": quantity, "warehouse_id": warehouse_id, "current_stock": _INVENTORY_DATA[product_id]["stock"][warehouse_id]} return {"error": "Unsupported API operation", "status_code": 400} def query_product_inventory(product_id: str, warehouse_id: Optional[str] = None) -> Dict[str, Any]: """Queries the private inventory management API for product stock.""" # In a real scenario, use requests directly # response = requests.get(f"{API_BASE_URL}/inventory/{product_id}", params={"warehouse_id": warehouse_id}) # response.raise_for_status() # return response.json() return _mock_api_call("GET", "/inventory/query", {"product_id": product_id, "warehouse_id": warehouse_id}) def add_product_inventory(product_id: str, quantity: int, warehouse_id: str) -> Dict[str, Any]: """Adds stock to a product in a specific warehouse via the private inventory API.""" # In a real scenario, use requests directly # payload = {"quantity": quantity, "warehouse_id": warehouse_id} # response = requests.post(f"{API_BASE_URL}/inventory/{product_id}/add", json=payload) # response.raise_for_status() # return response.json() return _mock_api_call("POST", "/inventory/add", {"product_id": product_id, "quantity": quantity, "warehouse_id": warehouse_id}) ``` That mock is a demo, and the gap between it and production is where the real work sits. Against a real database: connection pooling, parameter binding to prevent SQL injection, and queries you have actually looked at. Against a real API: error handling, timeouts, retries with exponential backoff, and caching for data that doesn't change. As a competitive programmer, every layer counts, and this is the layer that counts most. The agent's job is to call the right tool with the right arguments. The tool itself has to be tuned. **The tell:** if you can't import your backend module and call the function from a plain Python REPL with LangChain uninstalled, the framework has leaked into your logic, and swapping orchestration layers later means a rewrite. --- ### Step 3 — Wrap it, and spend your effort on the description `StructuredTool.from_function` joins the schema to the function. Two fields do the work, and they answer different questions. ```python # tools/__init__.py from langchain.tools import StructuredTool # Import schemas and backend functions from .schemas import ( StudentGradeInput, UpdateStudentGradeInput, InventoryQueryInput, AddInventoryInput ) from .backend import ( get_student_grade, update_student_grade, query_product_inventory, add_product_inventory ) # --- Student Grade Tools --- get_student_grade_tool = StructuredTool.from_function( func=get_student_grade, name="GetStudentGrade", description="Useful for retrieving a student's grade(s) from the school database. " "Can fetch all grades for a student or a specific course grade.", args_schema=StudentGradeInput ) update_student_grade_tool = StructuredTool.from_function( func=update_student_grade, name="UpdateStudentGrade", description="Useful for updating an existing student's grade in a specific course in the school database. " "Requires student ID, course name, and the new numerical grade.", args_schema=UpdateStudentGradeInput ) # --- Inventory Management Tools --- query_inventory_tool = StructuredTool.from_function( func=query_product_inventory, name="QueryProductInventory", description="Useful for querying the current stock level of a product " "in the private inventory management system. Can specify a particular warehouse.", args_schema=InventoryQueryInput ) add_inventory_tool = StructuredTool.from_function( func=add_product_inventory, name="AddProductInventory", description="Useful for adding stock to a product in a specific warehouse " "within the private inventory management system. Requires product ID, quantity, and warehouse ID.", args_schema=AddInventoryInput ) # List of all custom tools ALL_CUSTOM_TOOLS = [ get_student_grade_tool, update_student_grade_tool, query_inventory_tool, add_inventory_tool, ] ``` `description` decides **when** the tool gets used. It's the natural-language instruction the model reads while choosing, so it has to be specific enough to separate this tool from its neighbours. `args_schema` decides **how** it gets used, and Pydantic enforces that part for you. **The tell:** read your four descriptions back to back and ask which one answers the sentence "how many laptops are in W01". If two of them are plausible answers, the model is choosing between them on vibes. --- ### Step 4 — Wire the agent and let it orchestrate Only now does the framework appear. The agent's whole job is translation. ```python # main_agent.py import os from dotenv import load_dotenv from langchain.agents import AgentExecutor, create_openai_functions_agent from langchain_openai import ChatOpenAI from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder from tools import ALL_CUSTOM_TOOLS # Our custom tools load_dotenv() # Load OpenAI API key from .env # 1. Initialize the Language Model # Use a powerful, function-calling capable model llm = ChatOpenAI( model="gpt-4-turbo-preview", # Or "gpt-3.5-turbo" for cost-effectiveness temperature=0, api_key=os.getenv("OPENAI_API_KEY") ) # 2. Define the Agent Prompt # This is a standard prompt for OpenAI Functions Agent prompt = ChatPromptTemplate.from_messages([ ("system", "You are a helpful AI assistant with access to school database and inventory management systems. " "You can retrieve and update student grades, and manage product inventory."), MessagesPlaceholder(variable_name="chat_history"), ("human", "{input}"), MessagesPlaceholder(variable_name="agent_scratchpad"), ]) # 3. Create the Agent # This agent type is specifically designed for models that support function calling. agent = create_openai_functions_agent(llm, ALL_CUSTOM_TOOLS, prompt) # 4. Create the Agent Executor agent_executor = AgentExecutor(agent=agent, tools=ALL_CUSTOM_TOOLS, verbose=True) # --- Demonstrate Agent Capabilities --- async def run_agent_queries(): print("--- Query 1: Get Alice's Math grade ---") result = await agent_executor.invoke({"input": "What is Alice's grade in Math?", "chat_history": []}) print(f"Agent Response: {result['output']}\n") print("--- Query 2: Get Bob's all grades ---") result = await agent_executor.invoke({"input": "Tell me all grades for student Bob.", "chat_history": []}) print(f"Agent Response: {result['output']}\n") print("--- Query 3: Update Alice's Physics grade ---") result = await agent_executor.invoke({"input": "Alice's Physics grade needs to be updated to 90.0.", "chat_history": []}) print(f"Agent Response: {result['output']}\n") print("--- Query 4: Verify Alice's Physics grade after update ---") result = await agent_executor.invoke({"input": "What is Alice's Physics grade now?", "chat_history": []}) print(f"Agent Response: {result['output']}\n") print("--- Query 5: Check Laptop inventory in Warehouse W01 ---") result = await agent_executor.invoke({"input": "How many Laptops do we have in Warehouse W01?", "chat_history": []}) print(f"Agent Response: {result['output']}\n") print("--- Query 6: Add Mouse inventory ---") result = await agent_executor.invoke({"input": "Add 10 more Mice to Warehouse W02.", "chat_history": []}) print(f"Agent Response: {result['output']}\n") print("--- Query 7: Check Mouse inventory after addition ---") result = await agent_executor.invoke({"input": "What's the current stock of Mice in Warehouse W02?", "chat_history": []}) print(f"Agent Response: {result['output']}\n") if __name__ == "__main__": import asyncio asyncio.run(run_agent_queries()) ``` Run it with `verbose=True` and the executor prints its own control flow: read the query, pick the tool, format the arguments from the Pydantic schema, execute, and feed the result back through `agent_scratchpad` until there's an answer to return. ```mermaid flowchart LR Q[query] --> A[agent] A --> D{tool needed} D -->|yes| T[structured tool] T --> S[scratchpad] S --> A D -->|no| O[answer] ``` `ChatOpenAI` on a function-calling model handles the translation from prompt to call. Everything downstream of the diamond is code you wrote. **The tell:** in the verbose trace, you should be able to point at the line where the tool was chosen and the line where the arguments were filled. Those are two different failure modes, and if you can't separate them you'll fix the wrong one. --- ### What breaks it **The mock is not the backend.** The in-memory dict never times out, never rate-limits, and never returns a partial row. Every failure the mock cannot produce is a failure your tool has no handling for, and the agent will surface it as a confused answer rather than an error. **Two descriptions that overlap.** The model picks a tool by reading prose. `QueryProductInventory` and `AddProductInventory` are far enough apart to be safe. Add a third inventory tool with a fuzzy description and you have introduced a coin flip into a system you thought was deterministic. **The tool creates what it was asked to look up.** `_mock_api_call` auto-creates an unknown `product_id` on POST rather than rejecting it. A hallucinated ID stops being a caught error and becomes a real record. When the model supplies the arguments, "be lenient with input" is a different decision than it was for human callers. --- ### When a custom tool is the wrong choice Custom tools cost you a schema, a function, a description, and a test surface for each one. That is not always worth paying. - **The capability is genuinely generic.** Web search and arithmetic are what the pre-built tools are for. Rewriting `Calculator` to feel bespoke is work with no output. - **The call really is one string.** `StructuredTool` and Pydantic buy you validation across multiple or complex arguments. With a single string input, plain `Tool` is the simpler construct and the schema is ceremony. - **You only need one orchestration layer, and LangChain isn't it.** LangChain is verbose, and it's a wrapper. If a raw OpenAI function-calling loop covers your case, write the backend functions and skip the framework. The tool logic is the same either way, which is the point of keeping it framework-agnostic. - **Your model doesn't do function calling.** `create_openai_functions_agent` is built for models that support it. Without that, the structured contract has nothing on the other end to honour it, and you are back to parsing text. --- ### The shift The framework is not where the work is. LangChain is an orchestration layer, and the value lives in the four functions underneath it that would run perfectly well without it. - **Write the contract before the logic.** A Pydantic schema is a machine-readable API. A docstring the model never sees is not. - **Keep the backend framework-agnostic.** You should be able to swap LangChain for a raw function-calling loop without touching a tool. - **Specialize deliberately.** Generic tools give generic results, and the Mixture-of-Experts analogy is a working blueprint rather than a metaphor. Look at the tools your agent has today and ask which of them the model has to guess its way into. That list is your work queue. --- Source: https://himanshuat.com/blogs/crafting-custom-tools-for-langchain-agents ======================================================================== --- title: "Database agents that know when they're wrong" description: "A SQL agent that errors is a nuisance. A SQL agent that returns a clean, confident, wrong number is a liability. Notes from building the scaffolding around text-to-SQL and document-store querying, and on why the benchmarks everyone quotes are partly measuring annotation noise." date: "January 20, 2026" url: "https://himanshuat.com/blogs/db-agents-that-know-when-theyre-wrong" --- # Database agents that know when they're wrong The query didn't error. It ran clean, returned a number, and the number was wrong because a join had fanned out and silently multiplied a SUM. Nothing in the stack complained. The figure went into a report and got acted on. If an error like that is caught at all, it is caught weeks later, after the decision it informed has already been made. I built the query layer for a system whose numbers were read and used without anyone re-deriving them by hand. That constraint, rather than any interest of mine in benchmarks, drove everything below. > The job of a database agent is not to be accurate. It is to make wrongness loud, because a system that fails loudly belongs to a different category from one that fails quietly. ### Chapter 0 — What silent wrongness is A query fails in one of three ways. It can fail to parse or execute, which is free to detect. It can fail to answer the question that was asked, which is correct SQL for a different question. Or it can answer the right question with a result set that is quietly the wrong shape. Only the first one announces itself. The other two arrive as a clean table of numbers with no error attached, which is why the work here is scaffolding rather than prompting. The model is not the part of the system that notices. Both targets in this post are the same problem in different clothes: Postgres, and a document store. --- ### Step 1 — Read the leaderboard as a noise measurement BIRD is the benchmark I keep returning to, because its authors bothered to measure what a human does on the same questions. The human baseline is **92.96% execution accuracy**, set by data engineers and database students, with the figure dated December 16 2025. The top of the leaderboard reaches roughly 82% on test. | BIRD system | dev | test | |---|---|---| | Human baseline | — | 92.96 | | AskData + GPT-4o | 77.64 | 81.95 | | Agentar-Scale-SQL | 74.90 | 81.67 | Every one of those numbers is produced under what BIRD calls oracle knowledge, meaning an external evidence string is supplied alongside the question. That is a generous setting compared to production. The benchmark itself is 95 databases totalling 33.4 GB, with 12,751 question-SQL pairs. The roughly eleven point gap to the human baseline has persisted, and the closing that did happen came out of multi-stage pipelines rather than raw model capability. Spider 2.0 paints a harsher picture. I cite only the paper baselines. | Spider 2.0 baseline | Snow | Lite | |---|---|---| | Spider-Agent + o1-preview | 23.58 | 23.03 | | Spider-Agent + GPT-4o-2024-11-20 | 12.98 | 13.16 | The most clarifying line comes from the project site: "GPT-4o achieves only 10.1% success on Spider 2.0, compared to 86.6% on Spider 1.0." That sentence quantifies how much of Spider 1.0 was benchmark rather than problem. I deliberately don't cite the top of the public Spider 2.0 leaderboard. Snow shows 96.70 while Lite tops out at 76.23 on the same 547 examples under different execution settings. Those entries are self-reported, and a configuration scoring 96.70 on a set of problems where a sibling configuration scores 76.23 on that identical set is not describing something physically sensible. Then there is what the rankings are made of. The paper that changed how I read all of this is arXiv 2601.08778, which also appeared at CIDR 2026. The authors re-annotated benchmark data and found an annotation error rate of **52.8% on BIRD Mini-Dev** and **62.8% on Spider 2.0-Snow**. Re-evaluating all sixteen open-source BIRD leaderboard agents against corrected data moves results by anywhere from -7% to +31%, and moves rankings by -9 to +9 positions. The number that matters most is the correlation. Spearman correlation against the full set is r=0.85 on the uncorrected subset and **r=0.32 on the corrected subset**. > The ordering you see on a text-to-SQL leaderboard tracks annotation noise about as much as it tracks capability. The error categories are mundane: timestamp casting, join semantics, output-format ambiguity. They are the same things that go wrong in real analytics work. *One honest caveat.* The arXiv and CIDR versions of that work disagree with each other. The arXiv abstract gives 62.8% for Spider 2.0-Snow; the CIDR version gives 66.1%. I cite the arXiv figure and flag that the two versions differ, because quoting one without acknowledging the other would be presenting more precision than the source supports. Execution accuracy has its own problems underneath all of that, though here I'm on softer ground. The ETM paper in MDPI Future Internet 17(8):325 is reported to find false-positive rates up to 23.0% and false-negative rates of 28.9% for execution accuracy and exact set match. *I have not read its methodology, so I treat those as an indication of magnitude rather than as measurements I'd defend.* The mechanism needs no figure: a false positive means two queries returned identical result sets on one test database and were scored equivalent when they are not. The lineage of fixes runs from test-suite accuracy (arXiv 2010.02840), the official Spider metric since 2020, through ETM's tree matching to LLM-based equivalence judging. **The tell:** if you picked your architecture by reading down a leaderboard, you picked it on r=0.32. Read the ablations inside the papers instead, because those compare a system against itself. --- ### Step 2 — Spend on scaffolding in order of value per unit of engineering The systems that do well are pipelines, and their published ablations tell you which pieces earn their cost. CHESS runs four agents: Information Retriever, Schema Selector, Candidate Generator, Unit Tester. It reaches 71.10% on BIRD test, within 2% of the leading proprietary method at the time, with roughly 83% fewer LLM calls. The Schema Selector alone accounts for around +2% accuracy and a 5x token reduction, which is the clearest evidence I know that narrowing the schema before generation buys more per unit of effort than anything else in the pipeline. *The rest of what I read, I'm citing secondhand from reported figures rather than from papers I've worked through, so I use them for direction and not for decimals.* XiYan-SQL is reported at 75.63% on BIRD and 89.65% on Spider test, with an ablation showing that replacing its trained candidate selector with plain self-consistency costs about 3 points. If that holds, generate N and vote is the weak version of candidate selection, which matches what I saw without giving me a number I'd quote. MAC-SQL, reported at 59.59 EX on BIRD test, uses a Selector, Decomposer and Refiner. What I took from it is architectural rather than numerical: its Refiner checks syntax, execution feasibility, and empty result sets. Treating an empty result as a suspicious signal rather than as an answer is close to free and catches a whole family of wrong filters. DAIL-SQL is reported at 86.6% on Spider, and its stated core finding changed how I do few-shot selection. Models learn the mapping between a question and a SQL *skeleton*, so examples should be selected by skeleton similarity rather than surface text similarity. PremSQL takes the simplest version of repair, execution-guided decoding, where you append the database error to the context and regenerate, with a reported default cap of five trials. Ordered by value per unit of engineering: 1. Parse and syntax repair, because it is nearly free. 2. Execution errors fed back and regenerated. 3. Empty results treated as suspicious. 4. Multiple candidates with a real selector. 5. LLM-generated natural-language unit tests over the returned rows. The first is where a parser earns its place. ```python # sql_guard.py """AST-level checks over model-generated SQL. sqlglot is a parser and a transpiler. It is not a validator, and nothing here is sqlglot promising safety: every rule below is one I wrote on top of the AST, and a query that passes all of them can still be semantically wrong. """ from sqlglot import exp, parse_one from sqlglot.errors import ParseError ALLOWED_TABLES = { "analytics.v_orders", "analytics.v_customer", "analytics.v_product", } FORBIDDEN = ( exp.Insert, exp.Update, exp.Delete, exp.Drop, exp.Create, exp.Alter, exp.Merge, exp.Grant, exp.Command, ) MAX_ROWS = 5_000 class Rejected(Exception): """Carries a message the agent is allowed to read and retry against.""" def guard(sql: str, dialect: str = "postgres") -> str: try: tree = parse_one(sql, read=dialect) except ParseError as err: raise Rejected(f"unparseable: {err}") from err if not isinstance(tree, exp.Select): raise Rejected(f"root is {type(tree).__name__}, expected Select") for node in tree.walk(): if isinstance(node, FORBIDDEN): raise Rejected(f"forbidden node: {type(node).__name__}") # CTE names are legal table references that will never be in the allowlist. cte_names = {c.alias_or_name for c in tree.find_all(exp.CTE)} for table in tree.find_all(exp.Table): name = table.sql(dialect=dialect, identify=False) if name in cte_names: continue if name not in ALLOWED_TABLES: raise Rejected(f"table not in allowlist: {name}") limit = tree.args.get("limit") if limit is None or int(limit.expression.this) > MAX_ROWS: tree = tree.limit(MAX_ROWS) return tree.sql(dialect=dialect) ``` The rejection message is the point of that design. Every `Rejected` is text the model can read and correct against, which is the same loop PremSQL runs on database errors. The alternative you see everywhere is a regex denylist over `INSERT|UPDATE|DROP` and friends, and it is trivially defeated. Comments split keywords, CTEs hide the shape of the statement, and string literals trip naive matchers while genuinely destructive statements slip past differently naive ones. Walking the AST and allowlisting is the only version I'd defend in a review. **The tell:** if your validator can be defeated by a comment in the middle of a keyword, you have a string matcher, not a guard. --- ### Step 3 — Move enforcement out of the prompt and into the database Even a good validator is a piece of application code that someone will eventually bypass, so the real enforcement belongs where the query runs. ```sql -- roles.sql CREATE ROLE llm_agent LOGIN PASSWORD :'agent_password'; ALTER ROLE llm_agent SET default_transaction_read_only = on; ALTER ROLE llm_agent SET statement_timeout = '10s'; ALTER ROLE llm_agent SET idle_in_transaction_session_timeout = '15s'; ALTER ROLE llm_agent SET lock_timeout = '2s'; ALTER ROLE llm_agent CONNECTION LIMIT 8; REVOKE ALL ON SCHEMA public FROM llm_agent; GRANT USAGE ON SCHEMA analytics TO llm_agent; -- An explicit view list, never a blanket GRANT on the schema. GRANT SELECT ON analytics.v_orders TO llm_agent; GRANT SELECT ON analytics.v_customer TO llm_agent; GRANT SELECT ON analytics.v_product TO llm_agent; -- With the role configured this way, an INSERT emitted by the model fails -- with SQLSTATE 25006, "cannot execute INSERT in a read-only transaction", -- no matter how the request was phrased or what the system prompt said. -- Cost gate: run this on the same connection before the real query and -- reject when estimated rows or total cost exceed budget. Catches the -- accidental cross join before it executes rather than after. EXPLAIN (FORMAT JSON) SELECT c.region, sum(o.amount) FROM analytics.v_orders o JOIN analytics.v_customer c ON c.customer_id = o.customer_id GROUP BY 1; ``` Point the agent at a read replica rather than the primary, cap rows on the fetch path, and size the pool so a runaway agent can't exhaust it. The worst outcome of a bad generation is then a wasted ten seconds. The `EXPLAIN` gate has caught more genuine mistakes for me than any prompt instruction, because a missing join predicate shows up as an absurd row estimate long before it shows up as a wrong answer. It is the cheapest place to catch the fan-out that opened this post. On reference implementations: LangChain now flags `AgentExecutor` as legacy, so `create_sql_agent` is not the thing to build a 2026 system on, whatever the tutorials still say. **The tell:** ask what happens if the model emits a `DROP` and every layer of your application code is bypassed. If the answer isn't an SQLSTATE, your boundary is a paragraph of English. --- ### Step 4 — Classify the ambiguity before you generate anything This is the part I underestimated. Most wrong answers weren't failures of SQL generation. They were correct SQL for a different question than the one asked. Metric definitions come first. "Active user" and "revenue" are not English words in a warehouse, they are decisions someone made, and the model will invent a plausible version of that decision. Grain and deduplication come second and are the most dangerous. `COUNT(*)` against `COUNT(DISTINCT ...)` diverges quietly, and a join that fans out multiplies a SUM without any signal that anything happened. That single pattern accounts for more silent wrong answers in analytics than everything else combined. Then time boundaries, where inclusive versus exclusive endpoints, timezone handling and fiscal versus calendar periods all produce answers that look right. Then soft-delete and status filters the model has no way to know exist, because nothing in the schema announces that `orders` holds rows everyone already knows to exclude. Then currency and unit mismatches. Then NULL semantics, whose sharpest edge is `NOT IN` against a subquery containing a NULL, which returns zero rows rather than the answer you expected. *That list is mine, assembled from wrong answers I had to explain after the fact.* The nearest published taxonomy is AMBROSIA (NeurIPS 2024 Datasets and Benchmarks, arXiv 2406.19073), which I have only secondhand. It is reported to formalise three ambiguity types, scope, attachment and vagueness, and to find that the ambiguity persists even when the database is provided, with Ambrosia+ adding 2,535 unanswerable questions across 6,777 examples in 16 domains. If that second finding holds it is the important one, because it means this is not something you fix by pasting more schema into context. BIRD-Interact, an ICLR 2026 oral, benchmarks clarifying-question behaviour itself, which suggests the field now treats asking as a first-class action. Which means the output type has to admit it. ```python # contract.py from dataclasses import dataclass from typing import Literal Outcome = Literal["answer", "clarification_request", "refusal"] @dataclass(frozen=True) class Answer: kind: Outcome = "answer" sql: str = "" rows: list[dict] = None row_count: int = 0 checks_passed: list[str] = None @dataclass(frozen=True) class ClarificationRequest: kind: Outcome = "clarification_request" question: str = "" # "revenue gross or net of refunds?" ambiguity_class: str = "" # metric | grain | time | filter | unit | null candidates: list[str] = None # the readings the agent found plausible @dataclass(frozen=True) class Refusal: kind: Outcome = "refusal" reason: str = "" # no schema path, denied table, budget exceeded attempted_sql: str | None = None Result = Answer | ClarificationRequest | Refusal ``` Declaring the union is the easy part. The discipline is making sure the second and third arms are reachable often enough to be real, which means tracking their rate. **The tell:** an agent that has never returned a refusal in a month of live traffic is not well calibrated. Its refusal path is dead code. --- ### Step 5 — Screen the rows before a human reads them A handful of plausibility assertions catch a surprising share of the silent cases. ```python # plausibility.py from dataclasses import dataclass @dataclass class Suspicion: check: str detail: str def screen(rows, table_row_count, expected, controls) -> list[Suspicion]: """Cheap post-execution screening. None of these prove correctness; each one flags a shape that is usually a bug rather than a finding.""" out: list[Suspicion] = [] n = len(rows) if n == 0: out.append(Suspicion("empty", "zero rows: verify filters, not an answer")) if n == table_row_count: out.append(Suspicion("unfiltered", f"row count equals full table ({n})")) for row in rows: for col, val in row.items(): if col.endswith(("_pct", "_rate", "_share")) and val is not None: if not 0 <= float(val) <= 100: out.append(Suspicion("range", f"{col}={val} outside [0,100]")) # A fanned-out join shows up here: the parts sum past the known whole. for col, total in controls.items(): observed = sum(r[col] for r in rows if r.get(col) is not None) if observed > total * 1.001: out.append( Suspicion("control_total", f"sum({col})={observed} exceeds {total}") ) if expected.get("shape") == "scalar" and n > 1: out.append(Suspicion("cardinality", f"{n} rows for a scalar question")) return out ``` The strongest check isn't in that function, because it costs a second query: generate two structurally different queries for the same question and compare results. When they disagree you have found a real ambiguity, and the right output is a clarification request rather than a coin flip. Assembled, the pieces from Steps 2 through 5 are one path with two loops in it and three exits. ```mermaid flowchart TD Q[question] --> A{ambiguous?} A -->|yes| C[clarify] A -->|no| G[generate sql] G --> V{ast guard} V -->|reject| G V -->|pass| E{explain gate} E -->|over budget| R[refuse] E -->|pass| X[execute] X -->|db error| G X -->|rows| S{screen} S -->|suspicious| C S -->|clean| ANS[answer] ``` The two edges pointing back into generation are the whole design. Everything that can be turned into a message the model reads and retries against, should be. **The tell:** if a zero-row result reaches the user as an answer rather than as a question, you have shipped the most common silent failure in the system. --- ### Step 6 — Stop treating a document store as SQL with different syntax I expected MongoDB to be roughly text-to-SQL with different syntax. It is harder for four reasons that don't overlap. There is no schema to link to. Schema linking is the single most productive technique in text-to-SQL, and in a document store it has no input, so you induce a schema by sampling documents and hope the sample was representative. Polymorphism makes that worse. The same field is a scalar in one document, an array in another, absent in a third, and any induced schema flattens that away. Denormalisation then destroys the join signal, because the foreign key graph is what tells you how entities relate, and in a document store the relationship is implicit in nesting. The fourth reason is the one I find most interesting. Aggregation pipelines are programs, not declarations. Stage ordering changes semantics, and no planner will normalise a mistake the way a SQL optimiser quietly rescues a badly ordered set of predicates. Two specific behaviours produce silent wrongness on their own. `$unwind` on a missing field drops documents without comment. And `{status: {$ne: "cancelled"}}` matches documents where `status` doesn't exist at all. Both return a smaller, clean-looking result set and no error. TEND (arXiv 2502.11201) is the benchmark for this, with 1,210 MongoDB-native tasks across 11 databases. Its citable finding is that models with strong NL2SQL performance degrade substantially on it. There is no per-model accuracy figure to quote, only that direction, and the direction matches what I saw well enough to distrust anyone reporting document-store agent quality by analogy to their SQL numbers. **The tell:** if your document-store agent's quality estimate came from your SQL numbers, you have no estimate. --- ### Semantic layers, and the caveat that beats the win The dbt Developer Blog published a piece titled "Semantic Layer vs. Text-to-SQL: 2026 Benchmark Update", run on the data.world ACME Insurance benchmark with a 15-table schema. | model | text-to-SQL | with semantic layer | |---|---|---| | Claude Sonnet 4.6 | 90.0 | 98.2 | | GPT-5.3-Codex | 84.1 | 100.0 | I find their caveats more valuable than their headline. The benchmark is 11 questions run 20 times each, for 220 queries total, and n=11 is a very small question set for an architectural conclusion. They loaded the entire schema as context and concede this "isn't practical for larger datasets", which quietly removes one of the harder parts of the real problem. Some questions required three additional dbt models before MetricFlow could answer them at all, so part of what improved was the data model rather than the query interface. And more reasoning effort bought latency without buying accuracy. The line I'd anchor on isn't the accuracy delta at all. > "Text-to-SQL returns plausible but incorrect answers; the Semantic Layer returns errors when unable to answer." That categorical difference is the real argument for a semantic layer, and it is the same argument as everything above it in this post. The counterargument deserves airing. MotherDuck has argued that BIRD's difficulty is largely an artifact of badly modelled schemas, which implies semantic-layer gains partly measure data-modelling effort rather than architecture. I think that's substantially right, and it relocates the case rather than weakening it: if the fix is modelling work, the semantic layer is where that work becomes durable. --- ### Six things to build this week Rebuilding this tomorrow, this is the order I'd do it in. 1. A read-only role with an explicit view allowlist and timeouts. 2. AST validation with an allowlist rather than a denylist. 3. An `EXPLAIN` budget gate on the same connection as the query. 4. Empty results routed to suspicion rather than to the user. 5. A three-armed output type with a monitored refusal rate. 6. An ambiguity classifier in front of generation. --- ### What breaks it **The refusal and clarification arms quietly die.** They are the two arms that carry the entire value of the design, and they are also the two that nobody notices going missing, because a system that always answers looks like a system that works. The rate is the monitor. Without it the union type is documentation. **Every check in here is a shape check.** The comment at the top of `sql_guard.py` is the honest version: a query that passes every rule can still be semantically wrong. Plausibility screening flags shapes that are usually bugs, not shapes that are provably bugs. None of it proves correctness, and a system built on it will still be wrong sometimes. The claim is only that it will be wrong loudly more often. **Component choice by leaderboard.** At r=0.32 on corrected data, and with rankings moving -9 to +9 positions, the ordering is not a design input. The self-comparisons inside the papers are: CHESS's Schema Selector ablation is worth more to you than CHESS's rank. And several of the figures I've quoted here, including XiYan-SQL, MAC-SQL, DAIL-SQL and ETM, are secondhand. Use them for direction and not for decimals. --- ### When a database agent is the wrong choice Every layer above costs engineering time, a second query, and a slower answer. Most of that cost only pays back under one specific condition. - **Someone re-derives the numbers by hand anyway.** The entire premise here is that figures get read and acted on without a check. If a human already checks every number, you are paying for a second checker. - **Nobody can settle the metric definitions.** The clarification arm only helps if there is someone who can answer "revenue gross or net of refunds?" and make it stick. Without that owner, the classifier just relocates the argument into your inbox. - **You can't get a restricted role or a replica.** Step 3 is where the guarantees actually live. If the only thing between the model and the primary is application code and a system prompt, you don't have a boundary, and no amount of AST validation creates one. - **The schema is the real problem.** If MotherDuck is right that difficulty is largely an artifact of badly modelled schemas, then a query agent over a badly modelled warehouse mostly produces plausible answers faster. Do the modelling work first. --- ### The shift I went in expecting the hard part to be generation and the interesting part to be model choice. Both were wrong. The wrong answers that cost me something were correct SQL for a question nobody had asked, and none of them would have been prevented by a better model. - **Design for loud failure, not for accuracy.** A system that returns errors when it can't answer is in a different category from one that returns plausible numbers. - **Put the guarantees where the model cannot reach them.** A read-only role and a statement timeout hold when the prompt, the validator, and your own application code have all been bypassed. - **Monitor the rate at which the agent asks and refuses.** It is the only number in the system that tells you the calibration is alive. Model choice moves things less than any of that. The benchmark numbers, once you account for the annotation noise inside them, move things less still. --- Source: https://himanshuat.com/blogs/db-agents-that-know-when-theyre-wrong ======================================================================== --- title: "Deterministic retrieval, and the parts of RAG that refuse to sit still" description: "I asked a resume search for candidates with five years of experience and got someone with three. Embeddings encode semantic proximity, not predicate satisfaction. A worked account of hybrid SQL plus vector retrieval, the post-filter trap in pgvector, and why determinism is a property of the whole stack rather than a temperature setting." date: "February 17, 2026" url: "https://himanshuat.com/blogs/deterministic-rag" --- # Deterministic retrieval, and the parts of RAG that refuse to sit still I asked a resume search system for candidates with at least five years of experience. It returned someone with three. The retrieval was working exactly as designed, which was the problem. A shortlist that silently includes the wrong people is also silently excluding the right ones, and the second failure is invisible from both ends. The candidate never learns they were dropped. The person reading the list is relying on it being complete, with no way to check that it is. Somebody missing a callback because of an index parameter is not a tuning issue. Unpicking it took me through pgvector's index internals and into a question I had not thought to ask: which parts of a retrieval stack are allowed to be approximate at all. > Determinism is not a temperature setting. It is a property of every layer in the stack, and the layer that broke mine was the index. ### Chapter 0 — Embeddings encode proximity, not predicate satisfaction "5 years of experience" and "3 years of experience" are near neighbours in embedding space. They share almost all of their tokens, they occupy the same semantic region, and by every measure a bi-encoder is trained to optimise they are similar. Under a filter they are opposites. The embedding is doing its job. The job is simply not the one I was asking it to do. The same failure recurs across a whole family of query terms. Negation gets flattened, because "no Python experience" sits close to "Python experience". Identifiers, part numbers and version strings get treated as approximately equal to their neighbours, so `v2.14.1` retrieves `v2.14.0` happily. Anything where the answer depends on an exact discrete match rather than a region of meaning is a bad fit for cosine distance. --- ### Step 1 — Put hard constraints in SQL and soft preferences in ranking The rule I now hold to without exception is that hard constraints belong in SQL predicates, soft preferences belong in ranking, and the two must never be mixed. > Once a requirement is genuinely binary, expressing it as a similarity threshold is a category error dressed up as a tuning problem. No amount of raising the threshold makes "at least five years" true of the three-year candidate. **The tell:** if raising a similarity threshold ever felt like it was fixing a filter, it wasn't. It was discarding good candidates alongside the bad one. --- ### Step 2 — Turn iterative scan on, then treat its budget as a failure mode Having decided to put hard constraints in SQL, I hit the second failure, which is subtler and which pgvector documents plainly in its own README: "With approximate indexes, queries with filtering can return less results since filtering is applied AFTER the index is scanned." Read that carefully, because it describes a silent failure. You ask for 10 results. The HNSW index returns its 40 nearest neighbours. Your `WHERE` clause eliminates 35 of them. You get 5 rows back, your application renders 5 cards, and nothing anywhere raises an error. The system did not tell you that it looked at a small, unrepresentative slice of the candidate set and that the answer you wanted was in the part it never examined. pgvector's fix is iterative scan, and it is worth knowing the exact knobs. ```sql -- retrieval_settings.sql -- Index build. m defaults to 16, ef_construction to 64; both are build-time -- and changing either means a rebuild. CREATE INDEX ON resumes USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64); -- Full-text side, same table. CREATE INDEX ON resumes USING gin (search_tsv); -- Session settings. ef_search defaults to 40 and is tunable per session with -- no rebuild, which makes it the only quality/latency dial worth exposing. SET hnsw.ef_search = 100; -- Without this, a restrictive WHERE clause silently truncates the result set. -- strict_order preserves exact distance ordering; relaxed_order is faster and -- may return results slightly out of order. SET hnsw.iterative_scan = strict_order; -- Bounds on how hard iterative scan is allowed to work before giving up. SET hnsw.max_scan_tuples = 20000; -- default 20000 SET hnsw.scan_mem_multiplier = 1; -- default 1 -- IVFFlat, if you use it instead: lists = rows/1000 up to 1M rows, and -- sqrt(rows) above that. probes defaults to 1, which is almost never enough. -- CREATE INDEX ON resumes USING ivfflat (embedding vector_cosine_ops) -- WITH (lists = 1000); -- SET ivfflat.probes = 10; ``` The bounds matter as much as the switch. `max_scan_tuples` at its default of 20,000 means iterative scan is not a promise of completeness. It is a budget, and when the budget runs out you are back to a truncated result set with no error. What changes is that you now know the shape of the failure and can log when the budget was exhausted, which is the whole game. The check I added around all of this is embarrassingly simple and has been worth more than the tuning. If the query asked for k results and fewer than k came back, the response carries a flag saying the result set may be truncated by filtering rather than exhausted by the corpus. Those are very different statements, and the raw row count cannot distinguish them. Surfacing the difference costs one boolean and turns a silent recall loss into something a caller can act on, which is the same move as treating an empty SQL result as suspicious rather than as an answer. **The tell:** ask for 10 and count what comes back. If fewer than 10 arrive and nothing in the response says why, you cannot tell a truncated result set from an exhausted corpus. --- ### Why this is a bind, not an implementation gap I spent a while assuming pre-filtering was the obvious correct answer that pgvector had simply not implemented well. It isn't. | approach | recall | speed | failure mode | |---|---|---|---| | pre-filter | full, over the filtered subset | brute-force scan of whatever the filter selected | slow, but predictably slow | | post-filter | lost by a factor that depends on how correlated the filter is with the vector neighbourhood | fast, the index does the work | silent truncation, no error | There is no query-specific index for an arbitrary predicate, which is why pre-filtering degenerates into a scan. Neither side is a bug. The bind comes from three properties of ANN structures: they optimise for proximity rather than boolean predicate satisfaction, deleting nodes fragments the HNSW graph, and you cannot jointly index a vector and its metadata the way a relational engine indexes a composite key. There is active work on this. Curator (arXiv 2601.01291) and JAG (arXiv 2602.10258) both attack filtered ANN directly, and arXiv 2606.14193 makes the argument that existing filtered-ANN benchmarks are unrealistic enough that reported progress may not transfer. I read that last one as a caution against assuming the problem is about to be solved out from under you. --- ### Step 3 — Fuse ranks, not scores The other half of the design is combining lexical and vector retrieval, where the standard mistake is trying to add a BM25 score to a cosine distance. Those two quantities share no scale, and the relationship between them is not stable across queries, so any weighted sum you write is a set of magic constants that will drift. Reciprocal Rank Fusion sidesteps that. The score is a sum over rankers of `1 / (k + rank(d))`, with k conventionally set to 60. The technique is usually attributed to Cormack, Clarke and Buettcher at SIGIR 2009, and *I'm taking both that attribution and that constant from secondary sources rather than from the paper itself.* Neither carries the argument, which stands on the mechanism: fusing ranks rather than scores means the only thing two retrievers need in common is an ordering, and orderings are always comparable. The shape of the query is a fan-out and a merge, with the same predicates on both arms. ```mermaid flowchart TD Q[query] --> L[lexical CTE] Q --> S[semantic CTE] P[hard predicates] --> L P --> S L --> F[RRF fuse] S --> F F --> R[ranked 20] ``` ```sql -- hybrid_search.sql -- One query, two retrievers, RRF combine. The metadata predicates are pushed -- into BOTH CTEs: filtering only one side reintroduces the post-filter trap -- through the back door. WITH params AS ( SELECT $1::vector AS q_vec, $2::text AS q_text, $3::int AS min_years, $4::text[] AS locations, 60::float AS k, 200::int AS depth ), lexical AS ( SELECT r.id, row_number() OVER ( ORDER BY ts_rank_cd(r.search_tsv, websearch_to_tsquery('english', p.q_text)) DESC ) AS rank FROM resumes r, params p WHERE r.search_tsv @@ websearch_to_tsquery('english', p.q_text) AND r.years_exp >= p.min_years AND r.location = ANY (p.locations) AND r.work_auth = true LIMIT (SELECT depth FROM params) ), semantic AS ( SELECT r.id, row_number() OVER (ORDER BY r.embedding <=> p.q_vec) AS rank FROM resumes r, params p WHERE r.years_exp >= p.min_years AND r.location = ANY (p.locations) AND r.work_auth = true ORDER BY r.embedding <=> p.q_vec LIMIT (SELECT depth FROM params) ) SELECT COALESCE(l.id, s.id) AS id, COALESCE(1.0 / (p.k + l.rank), 0.0) + COALESCE(1.0 / (p.k + s.rank), 0.0) AS rrf_score, l.rank AS lexical_rank, s.rank AS semantic_rank FROM lexical l FULL OUTER JOIN semantic s ON s.id = l.id CROSS JOIN params p ORDER BY rrf_score DESC LIMIT 20; ``` Two honest notes about that query. `ts_rank_cd` is cover density ranking, which is a reasonable lexical signal but is not BM25; if you want real BM25 inside Postgres then ParadeDB's `pg_search` is the thing that provides it. And returning both component ranks alongside the fused score is not decoration. When a result looks wrong, the first question is always whether it came from the lexical side, the semantic side, or scraped in from both, and you cannot answer that after the fact if you threw the ranks away. For the shape of the result, BEIR (NeurIPS 2021 Datasets and Benchmarks) is the citable source: dense retrievers that win in-domain lose out-of-domain, and BM25 remains a highly competitive zero-shot baseline across its 18 datasets. *I'm deliberately not attaching a decimal to that, in either direction.* The figures people quote for how far dense beats BM25, or what hybrid adds on top, are aggregates I cannot trace back to a paper, and the ones for 2026 hybrid stacks live in blog leaderboards rather than anywhere I can check. Shape, not measurement. **The tell:** if a result looks wrong and you cannot say which retriever put it there, the ranks got thrown away somewhere upstream. --- ### Step 4 — Give the planner a third answer The pipeline that came out of all this parses each resume, extracts typed columns, and embeds the free text separately, so the structured and unstructured parts of a document are queried by the mechanisms suited to each. ```python # query_plan.py from dataclasses import dataclass, field from typing import Literal Outcome = Literal["answer", "clarification_request", "refusal"] @dataclass class HardConstraints: """Every field here becomes a SQL predicate. Nothing here is negotiable, and nothing here is ever expressed as a similarity threshold.""" min_years_exp: int | None = None work_authorized: bool | None = None locations: list[str] = field(default_factory=list) required_skills: list[str] = field(default_factory=list) # join table, exact @dataclass class SoftPreferences: """Ranking signals only. These can move a candidate up or down; they can never remove one from the result set.""" domain_text: str = "" # "fintech backend, high write throughput" prefers_skills: list[str] = field(default_factory=list) seniority_hint: str | None = None @dataclass class Plan: kind: Outcome hard: HardConstraints | None = None soft: SoftPreferences | None = None question: str | None = None # set when kind == clarification_request reason: str | None = None # set when kind == refusal def plan_from_request(req: str) -> Plan: """Extraction step. If a stated requirement cannot be mapped onto a typed column, it does not silently become a ranking signal: it becomes a clarification request, because demoting a hard constraint to a soft one is how you end up returning a three-year candidate for a five-year search.""" ... ``` That last docstring is the design in one sentence. If a user asks for "senior engineers near the office" and neither seniority nor proximity is a typed column, the honest response is a question, not a ranked list that quietly reinterprets the request. Two published findings point the same way, and *I have both secondhand rather than from the papers.* arXiv 2602.18550 is reported to find that many models cannot consistently select the resume describing the more qualified candidate, and do not reliably abstain when candidates are equally qualified. That second half is the more damning of the two, because a system that expresses a preference between two equivalent candidates is generating a preference rather than detecting one. Separately, Wilson and Caliskan, published at AIES, are reported to find that LLM resume screening disadvantages Black-associated and female-associated names with all other content held identical. I'd have made the same design call without either result, because a ranker whose reasoning you cannot inspect is already a problem before you know what it is keying on. What the findings add is a specific account of what it might be keying on, and that makes this a correctness argument as much as an ethical one. A name carries no information about the question being asked, so a ranker that responds to one is responding to noise. Hard filters over typed columns plus an auditable ranking function is a design where you can inspect what moved a candidate up or down. An end-to-end model judgement is one where you cannot, and "we could not tell you why" is not an answer that survives being asked about a specific person. **The tell:** if the clarification arm has never fired, it is not a path. It is an unreachable branch, and every ambiguous request is being silently reinterpreted as a ranked list. --- ### Step 5 — Make retrieval the fixed point and let generation vary Here is where my mental model was most wrong. I had assumed retrieval was the deterministic half of the system and that any run-to-run variation had to be coming from the model. pgvector says otherwise, in one sentence in its README: "Unlike typical indexes, you will see different results for queries after adding an approximate index." That is the entire warning, and it is enough. The moment you build an approximate index, retrieval stops being a function of the query alone and becomes a function of the index's internal state. Identical queries can return different rows. Anything downstream that changes between runs may be changing because of that, with the model behaving identically. The generation side turns out not to be a fixed point either, and *here I'm reporting rather than verifying.* Thinking Machines Lab's "Defeating Nondeterminism in LLM Inference" reports that sampling 1,000 completions from Qwen3-235B-Instruct at temperature 0 produced 80 distinct outputs, with the first divergence at token 103, and that batch-invariant kernels restore bitwise-identical output across 1,000 runs on Qwen3-8B at roughly a 61.5% throughput cost, falling to about 34.35% once CUDA graphs are applied. I have not reproduced any of those figures. The mechanism they describe is the part I'd act on, because it can be reasoned about without trusting the numbers. Matmul, RMSNorm and attention kernels change their floating-point accumulation order depending on batch size, batch size depends on how many concurrent requests the server is handling, and floating-point addition is not associative. Your output therefore depends on other people's traffic. > A seed pins the sampler. The divergence sits upstream of sampling, in the kernels themselves. So I don't chase bitwise determinism end to end, which is expensive and mostly unnecessary. I make retrieval deterministic and let only generation vary, so that when an answer changes I know which layer changed. **The tell:** when an answer changes between two runs, can you name the layer that moved? If not, four layers are varying at once and every investigation starts from zero. --- ### Step 6 — Gate on recall, separately The metric that dominates everything downstream is recall@k, because the generator cannot cite what it never saw. Whatever your answer quality is, it is bounded above by whether the right document made it into the context window. | metric | fits when | in this system | |---|---|---| | recall@k | anything feeding a generator | the ceiling on every downstream number | | MRR | exactly one correct answer and its position matters | not the shape of resume search | | nDCG@k | graded relevance | candidates are better and worse, not right and wrong | The apparatus that made this tractable was small: a golden set of somewhere between 50 and 200 query-to-expected-document pairs drawn from the actual corpus rather than a public benchmark, run in CI on every pull request that touches retrieval, with one variable changed at a time. Public benchmarks tell you about the general shape of retrieval behaviour. They tell you nothing about whether your chunking change broke your corpus. Building that set is the tedious part and there is no way around it. I labelled mine by hand, one query at a time, writing down which documents genuinely should have surfaced before looking at what the system returned, because doing it in the other order produces a golden set that ratifies the behaviour you already have. A hundred honestly labelled pairs from your own corpus beat any quantity of borrowed relevance judgements, and the labelling surfaces query ambiguities no metric would show you. **The tell:** a change that improves nDCG while dropping recall@50 is a regression, and an aggregate score will hide it. If your CI cannot fail on that specific combination, it is not gating recall. --- ### Four things to pin this week Step 5 is an argument. These four commits are what it looks like in a repository, and none of them takes long. - **Pin the embedding model version and store it per row.** Re-embedding a corpus with a new model version silently invalidates every stored vector and absolutely nothing errors. - **Pin and log `ef_search` per query class.** Then a latency tuning change shows up in the audit trail rather than as mysterious quality drift. - **Use RRF.** Rank-based fusion means small score perturbations don't reorder results. - **Snapshot the corpus by version.** An eval run is only reproducible against the corpus it was scored on. --- ### What breaks it **Iterative scan is a budget, and budgets get spent.** `max_scan_tuples` defaults to 20,000. Past that, the scan gives up and you are back to a truncated result set with no error raised. The switch changes the shape of the failure, not whether it can happen. Log the exhaustion or you have moved the silence rather than removed it. **Filtering one arm of the hybrid query.** The predicates have to be in both CTEs. Push them into the lexical side alone and the semantic side reintroduces the post-filter trap through the back door, with the fused result looking perfectly healthy. **A golden set built after looking at the output.** Label first, then run. Reverse that order and the set encodes the behaviour you already have, which means every future change is measured against your current bugs. **Re-embedding without a version column.** Nothing errors. The vectors are the wrong shape of wrong, distances stay plausible, and the only symptom is quality that drifts for reasons nobody can reconstruct. --- ### When this is the wrong choice Most of this machinery exists to defend a boundary between hard constraints and soft ones. If your system doesn't have that boundary, you are paying for scaffolding that holds nothing up. - **Every requirement in your queries is a soft preference.** No years, no locations, no work authorisation. Then there are no SQL predicates to push down, no post-filter trap to fall into, and plain vector search is the right amount of machinery. - **Your corpus is small enough for exact search.** The approximate index is what makes retrieval non-deterministic and what creates the filtering trap in the first place. Without one, most of this post is describing a problem you don't have. - **You only have one retriever.** RRF fuses orderings. With a single ordering there is nothing to fuse, and the constant 60 is just a number you copied. - **Nobody will label the golden set.** Fifty to two hundred hand-labelled pairs is real work by a person who knows the domain. Without it you have retrieval you can tune but not evaluate, and tuning against an unmeasured target is how you ship a confident regression. - **You need bitwise-identical generation.** That is a kernel problem, not a retrieval design. Batch-invariant kernels cost roughly 61.5% of throughput, about 34.35% with CUDA graphs, and no amount of pinning `ef_search` gets you there. --- ### The shift Retrieval determinism is achievable and worth paying for. Generation determinism depends on kernel-level batch invariance and is mostly not worth chasing. I had those two backwards for months, and the cost of having them backwards was not a slow system. It was a three-year candidate in a five-year shortlist, and no mechanism anywhere in the stack that could have told me. - **Hard constraints are predicates.** The moment one becomes a threshold, it stops being a constraint. - **Count what came back.** Fewer rows than you asked for is a fact about your index, not about your corpus, until something proves otherwise. - **Pin one layer so the other is legible.** Retrieval fixed, generation free, and a changed answer names its own cause. The concrete next step: run one filtered query, ask for 10, and check whether your application can tell you why it got 5. > If the answer is no, you don't have a ranking problem. You have a system that cannot distinguish an empty corpus from an index that never looked. --- Source: https://himanshuat.com/blogs/deterministic-rag ======================================================================== --- title: "10 Essential Programming Concepts Every Developer Should Know" description: "The core programming concepts that come up in almost every codebase." date: "February 11, 2023" url: "https://himanshuat.com/blogs/essential-programming-concepts" --- # 10 Essential Programming Concepts Every Developer Should Know A lot of day-to-day development rests on a handful of core concepts. Knowing them well makes it easier to write clear code and reason about the systems you build. ![Programming](https://unsplash.com/photos/ZG3UT6_RN6c) _Photo by Mohammad Rahmani on Unsplash_ Here are ten worth knowing well. ## Variables and Data Types Storing and manipulating data is the foundation. You need to know the data types a language offers and how variables hold values. ## Control Structures Control structures are used to determine the flow of a program. They include decision-making statements, loops, and other constructs that allow developers to control the execution of their code. ## Functions and Methods Functions and methods are reusable blocks of code you can call from anywhere in a program. Knowing how to define them, call them, and pass parameters is key to keeping code modular. ## Object-Oriented Programming Object-oriented programming is a programming paradigm that uses objects to represent real-world entities. Developers must understand the concepts of classes, objects, inheritance, and polymorphism to write effective object-oriented code. ## Exception Handling Exception handling is how a program responds to runtime errors. Handling them well keeps a program from falling over on the first unexpected input. ## Arrays and Collections Arrays and collections hold multiple values at once. Knowing which structure fits a given job keeps your code efficient. ## Input and Output Input and output are how a program talks to the outside world. Reading from and writing to files, streams, and other sources is part of almost everything useful. ## Algorithms and Data Structures Algorithms and data structures are essential tools for solving complex problems. Developers must understand the different types of data structures and how to use them to create efficient algorithms. ## Debugging Debugging is finding and fixing errors. Good tools and a systematic approach get you there faster than guessing. ## Testing Testing is the process of validating that a program meets its requirements. Developers must understand how to write and execute tests to ensure that their code works as intended. None of this is advanced, but it's the stuff you reach for every day. Get comfortable with it and the harder problems become a lot more approachable. --- Source: https://himanshuat.com/blogs/essential-programming-concepts ======================================================================== --- title: "FID was the wrong metric for the problem we had" description: "Aggregate FID went 22.1 to 18.5 and I nearly read that as months of adapter work buying four points. Half the eval set was already close to solved and it dragged the average toward nothing. Splitting by garment category was the change that made the result visible." date: "August 26, 2025" url: "https://himanshuat.com/blogs/fid-was-the-wrong-metric" --- # FID was the wrong metric for the problem we had Aggregate FID went from 22.1 to 18.5. I read that as several months of expert adapter training, structural conditioning and merge work buying about four points, and I nearly carried that number into a roadmap conversation as evidence the work hadn't paid off. The number was correct. The way I was reading it was not, and fixing that took one afternoon. > Any aggregate over a mixed population is a weighted average with the weights hidden. If half that population is already solved, the solved half isn't neutral. It's suppressing your signal. ### Chapter 0 — What a single distributional score can see FID is a distributional distance. You embed the generated set, embed the reference set, and compare the two distributions. That single number is computed over everything you fed it, which means the composition of the eval set is baked into the result and invisible in the output. Our evaluation set is 500 images split evenly between two categories we called Western Casual and Global Traditional. The first is t-shirts and jeans. The second is sarees, kimonos, hanfus and structured couture. We ran three systems over it: SDXL with a ControlNet baseline, base Flux Fill with no adapters and no structural conditioning, and our pipeline. Aggregate FID went 28.4, then 22.1, then 18.5. On the Western half, base Flux Fill was already close to good. Very little headroom, so almost nothing we did could move it. The other half is where every failure we cared about lived. ```mermaid flowchart LR E[eval set] --> W[western half] E --> G[global half] W -->|no headroom| A[aggregate FID] G -->|all headroom| A A --> R[reads as modest] ``` Average a category with no headroom against a category with enormous headroom and you get a number that moves by roughly half of what actually happened. You cannot tell from the number which half it came from. There's a sharper version. If we had shipped an improvement that fixed Global Traditional completely and slightly regressed Western Casual, the aggregate could have stayed flat. > A metric that can stay flat while the thing you built the company around gets solved is not measuring your problem. --- ### Step 1 — Write the failure down before you pick the metric We built the try-on pipeline to fix something specific: models that render a saree as a flat texture stuck to a torso, or classify a kimono as a bathrobe. That is a precise failure statement. We had it in hand, and then reached for the metric the field reports rather than the one that could see it. The order was backwards. Choose the metric after you've written the failure down, because the failure statement is what tells you which axis the metric has to be sensitive to. **The tell:** if you can state your failure in one sentence and the metric you're about to report cannot distinguish that sentence being true from it being false, you picked the metric first. --- ### Step 2 — Split the eval set on the axis you claim to be improving The fix is dull. Report per category, on the axis you're actually trying to move. Reported that way, with SSIM against the reference garment and a human-judged occlusion score: | Method | FID | SSIM (Western) | SSIM (Global) | Occlusion accuracy | | --- | --- | --- | --- | --- | | SDXL + ControlNet | 28.4 | 0.82 | 0.45 | 62% | | Base Flux Fill | 22.1 | 0.91 | 0.58 | 70% | | Ours | 18.5 | 0.94 | 0.85 | 92% | Read across the rows instead of down the FID column. Western SSIM goes 0.82, 0.91, 0.94. That's a good baseline getting slightly better, and it confirms what we already believed: Western casual wear is close to solved by any competent diffusion inpainter. Global SSIM goes 0.45, 0.58, 0.85. That's a different graph entirely. The baseline fails outright, plain Flux fails more politely, and the adapters do the thing they were trained to do. The gap between 0.94 and 0.85 is worth sitting with too. Even after the work, structurally complex garments score below simple ones. We closed most of the gap and did not close it. The split makes that honest in a way the aggregate never could, because the aggregate had no way to express that a gap existed at all. **The tell:** if a total success on the thing you care about could leave your headline number flat, you're reading a weighted average of your problem and somebody else's. --- ### Step 3 — Make the harness refuse an unbalanced manifest The split only stays honest if the categories stay balanced. Here is the shape of the harness that produces it: ```python # eval/run_split_eval.py """Score a checkpoint on the VTON eval set, split by garment category. The aggregate FID is kept because it is what everyone else reports and dropping it makes comparison impossible. It is reported alongside the per category structural scores, never on its own. """ from __future__ import annotations import argparse import json from collections import defaultdict from pathlib import Path from skimage.metrics import structural_similarity from aurax.eval.fid import frechet_distance from aurax.eval.io import load_pair, load_manifest CATEGORIES = ("western_casual", "global_traditional") def ssim_by_category(manifest: list[dict], run_dir: Path) -> dict[str, float]: """Mean SSIM between generated and reference garment, per category.""" scores: dict[str, list[float]] = defaultdict(list) for item in manifest: generated, reference = load_pair(run_dir / item["id"], item["reference"]) score = structural_similarity( generated, reference, channel_axis=-1, data_range=1.0, ) scores[item["category"]].append(score) return {cat: sum(v) / len(v) for cat, v in scores.items()} def evaluate(manifest_path: Path, run_dir: Path) -> dict: """Return the full report: aggregate FID plus per category SSIM.""" manifest = load_manifest(manifest_path) counts = {c: sum(1 for i in manifest if i["category"] == c) for c in CATEGORIES} if len(set(counts.values())) != 1: raise ValueError(f"eval set is not balanced across categories: {counts}") report = { "n": len(manifest), "per_category_n": counts, "fid_aggregate": frechet_distance(manifest, run_dir), "ssim": ssim_by_category(manifest, run_dir), } print(json.dumps(report, indent=2)) return report if __name__ == "__main__": ap = argparse.ArgumentParser() ap.add_argument("--manifest", type=Path, default=Path("eval/manifest.json")) ap.add_argument("--run", type=Path, required=True) evaluate(ap.parse_args().manifest, ap.parse_args().run) ``` The balance check in the middle is there because an unbalanced eval set silently reweights the aggregate. Once you've been burned by a number that averaged away your result, you start asserting the things you assumed. If somebody adds forty more t-shirts to the manifest next quarter, the harness should refuse to run rather than quietly report a slightly different FID. **The tell:** if a teammate can grow one category of your eval set and every number still prints, your aggregate is being reweighted and nothing in your pipeline will say so. --- ### Step 4 — Write the rating protocol before you collect the human number The fourth column is the one I'd defend least, and it's also the one that predicted product acceptance best. Occlusion accuracy is the rate at which a rater judged that hand and body occlusions were preserved correctly: the garment tucking behind a hand resting on a hip instead of being painted over it. We called the failure the Amputated Hand. For anyone who has looked at a lot of try-on output, 62% to 70% to 92% is the single most convincing sequence in the table. It was collected as a blind A/B test. That is genuinely all my notes record. I don't have the rater count, whether they were internal, how many images each one saw, how disagreements were resolved, or whether the same people scored all three systems. *I believe the comparison is directionally right, because the effect is large and visible without training.* But I can't hand you a protocol, and a number without a protocol is not one you should port into your own comparison. I didn't write it down, and the number is permanently weaker for it. **The tell:** if you can't say who rated, how many of them, and how disagreement was resolved, you have a direction rather than a measurement, and it belongs in the post with the hedge attached. --- ### What breaks it **The number you never defined.** The other figure attached to this work is an 85% success rate on complex global garments against about 15% for baseline models. It's the headline claim and the one I'd caveat hardest, because success was never given a precise rubric. In practice it meant somebody looked at the output and decided whether it was usable in a catalogue, which is a real and commercially meaningful judgment and is not a measurement. Different people draw that line differently, and the line probably moved as we got used to better output. The ratio is large enough that I don't think a stricter rubric would flip the conclusion, but the honest statement is narrower than the headline: on the garments where baselines fall over, our pipeline usually produces something a person would ship, and baselines usually don't. **FID at this sample size travels badly.** 500 images is a small sample for a distributional metric, and FID is biased upward at small sample sizes in a way that depends on the sample size. My notes also don't record which feature extractor we used. That combination means the FID column is fine for ranking three systems evaluated identically by us, and close to useless against a number in someone else's paper. Treat every cross-paper FID comparison you read with the same suspicion. **The split only sees the axis you chose.** Two failure modes survive at 0.85 Global SSIM and neither is legible in any column above. Extreme poses still break. Anything acrobatic or unusually articulated, yoga positions being the case we hit most, and the warping goes wrong. That's a data problem rather than an architecture problem: the base model and both of our adapters are trained on people standing or sitting, so the model has no prior for the pose and improvises. Our eval set inherits the same bias, which is why a split by garment doesn't catch it. A split by pose would. Sheer fabrics come out opaque. Lace and chiffon need the skin tone underneath to blend with the fabric on top, and the model renders them as solid cloth. SSIM against a reference garment is comparatively forgiving here, because the structure is roughly right and only the transparency is wrong, so the metric under-penalises a failure a customer would reject immediately. > The axis you split on is a claim about which variation matters. Ours was right, and it was also incomplete. Every uncaught failure mode above is a category we didn't think to separate. --- ### When splitting by category is the wrong choice The split costs you sample size and a second column to defend. That isn't always a trade worth making. - **The population is genuinely uniform on your axis.** If no subset has meaningfully different headroom, the split just prints the aggregate twice with more noise in each cell. - **You need to sit next to someone else's published number.** We kept aggregate FID for exactly this reason. Dropping it makes comparison impossible, so the split is an addition, never a replacement. - **Your cells get too small to support the metric.** 500 images is already small for FID, and every split you add halves what each cell is computed over. Structural per-image scores like SSIM survive the shrink better than distributional ones do. - **You haven't decided what variation you're claiming to fix.** A split is a claim about which axis matters. Split on garment when the failure is pose and you'll get two clean columns that both miss the problem. --- ### The shift We had a precise failure statement in hand, in the form of sarees rendering as flat texture, and then chose the metric the field reports instead of the one that could see it. Everything else in this post is downstream of that ordering mistake. The other thing worth carrying: the occlusion column moved product decisions and the FID column didn't. A crude metric with a written protocol beats a sophisticated one without, because people trust the thing they can picture. - **Write the failure down first, then pick the metric that can see it.** Not the other way round. - **Report per category on the axis you claim to be improving.** Keep the aggregate next to it, never on its own. - **Write the rating protocol before you collect the data.** Numbers you can't define, define your ceiling. Being unable to state what success meant is why the 85% figure sits in a blog post with three sentences of hedging around it instead of in a paper. The next thing to fix sits upstream of all of it. The Global Traditional half is exactly where fine structure matters most, pleats and embroidery, and that's a resolution problem before it's a modelling problem. > Which means confronting how much of a full frame we can actually afford to run at native resolution. --- Source: https://himanshuat.com/blogs/fid-was-the-wrong-metric ======================================================================== --- title: "Fitting a saree into 24GB: the context window trick" description: "A saree needs pixel density to hold its pleats, and a full-frame high-resolution pass doesn't fit on a 24GB card. The fix was to stop treating the frame as the unit of inference and hand the model only the region that matters." date: "September 02, 2025" url: "https://himanshuat.com/blogs/fitting-a-saree-into-24gb" --- # Fitting a saree into 24GB: the context window trick A saree fails at low resolution in a specific way. The pallu stops reading as fabric folded over a shoulder and becomes a printed pattern on skin. The pleats at the waist collapse into a smear. You need pixel density on the garment for the model to have anything to work with. What you can't do is buy that density by running the whole frame at high resolution, because a full body shot at the resolution the pleats need will not fit on a 24GB card. > The frame is not the unit of inference. The garment is. Once you accept that, 24GB stops being the binding constraint. ### Chapter 0 — What the model is actually paying for Start with what a full body e-commerce shot contains. A person occupies the middle of a tall frame. The garment occupies maybe a third of that. The rest is backdrop, floor, headroom, and a lot of pixels the diffusion model is going to encode, denoise for every step, and decode, in order to reproduce something we already had and didn't want changed. At native inference resolution the model is doing useful work everywhere. Scale the frame up to the point where the pleats survive and most of that work is now being spent on a grey sweep. Memory in a diffusion transformer scales with the number of latent tokens, and the token count grows with area. **Doubling the linear resolution of the frame roughly quadruples what you need to hold.** That's how you get to a place where a 48GB L40S is comfortable and a consumer 24GB card is not, on an image that is mostly backdrop. We had to care about 24GB because that's what a lot of the hardware our work would eventually run on looks like. Designing the pipeline so it only ever ran on datacentre cards would have been a decision with consequences we couldn't take back later. --- ### Why naive tiling breaks pleats The standard answer to a big image is to cut it into tiles, run each, and stitch. That's the right answer for upscaling a photograph of a hillside and the wrong answer here. A pleat is a continuous structure. It starts at the waist, runs down, and its shadow, direction and spacing are consistent along its whole length because gravity is consistent. Cut that structure across two tiles and each tile is denoised by a model that cannot see the other half. The model has a prior about fabric and it will happily produce plausible pleats in both tiles. It has no mechanism at all for making the pleat in tile B continue the pleat in tile A at the same angle, spacing and phase. You get two convincing halves of two different garments meeting at a line. Overlap and blending reduces the visible discontinuity without fixing it. Blending averages two incompatible structures, and the average of two out-of-phase pleat fields is a mush that reads as a manufacturing defect. This is the same reason tiled upscaling is bad at text and faces. It fails on anything where the correct value of a pixel depends on content further away than the tile. The structures that matter most for the garments we cared about are exactly the long range ones. So tiling was out, and the question became how to give the model high pixel density over a continuous garment without asking it to hold the whole frame. The answer is to change what the model sees rather than how it processes what it sees. Crop to the garment plus context, work at native resolution on that crop, put the result back. ```mermaid flowchart LR I[source frame] --> W[roi window] M[mask] --> W W --> U[upscale] U --> P[inpaint] P --> D[downscale] D --> C{composite} I --> C M --> C ``` The two edges into `composite` are the whole difficulty, and they arrive last. --- ### Step 1 — Crop to the mask, then pad it Take the bounding box of the final mask and pad it. The padding factor matters more than it looks. Too tight and the model has no surrounding body to anchor the garment against, and it will render a garment that doesn't know where the shoulders are. **We use rho = 0.2.** ```python # pipeline/context_window.py """Crop, upscale, inpaint and composite: the context window mechanism. The unit of inference is the garment region, not the frame. Everything here exists to make the model see a high pixel density crop while the delivered asset stays full resolution. """ from __future__ import annotations from dataclasses import dataclass import numpy as np NATIVE_RESOLUTIONS = ((1024, 1024), (1344, 1344)) ROI_PADDING = 0.2 @dataclass(frozen=True) class Window: """A crop rectangle in the coordinate space of the source image.""" x0: int y0: int x1: int y1: int @property def size(self) -> tuple[int, int]: return self.x1 - self.x0, self.y1 - self.y0 def roi_from_mask(mask: np.ndarray, rho: float = ROI_PADDING) -> Window: """Bounding box of the mask, padded by rho and clamped to the frame. rho is a fraction of the box dimension, applied on every side, so a tall narrow mask gets proportionally more horizontal context than a fixed pixel pad would give it. """ ys, xs = np.nonzero(mask) if ys.size == 0: raise ValueError("empty mask: nothing to inpaint") y0, y1 = int(ys.min()), int(ys.max()) x0, x1 = int(xs.min()), int(xs.max()) pad_y = int((y1 - y0) * rho) pad_x = int((x1 - x0) * rho) h, w = mask.shape[:2] window = Window( x0=max(0, x0 - pad_x), y0=max(0, y0 - pad_y), x1=min(w, x1 + pad_x), y1=min(h, y1 + pad_y), ) print(f"mask box: {(x0, y0, x1, y1)}") print(f"window: {(window.x0, window.y0, window.x1, window.y1)} size={window.size}") return window def pick_native(window: Window) -> tuple[int, int]: """Choose the inference resolution whose aspect is closest to the crop.""" crop_w, crop_h = window.size crop_aspect = crop_w / crop_h return min( NATIVE_RESOLUTIONS, key=lambda r: abs((r[0] / r[1]) - crop_aspect), ) ``` The padding is a fraction of the box rather than a fixed pixel count on purpose. A saree mask is tall and narrow. A fixed pad would give it a sliver of context on the sides, which is where it needs the most, because the model has to place the drape relative to a torso it can see. **The tell:** if the output garment doesn't know where the shoulders are, rho is too small. The model is being asked to drape a body it can't see. --- ### Step 2 — Upscale the crop to native resolution The crop gets resized up to whatever the model actually wants, which for us is 1024x1024 or 1344x1344. `pick_native` chooses between them by comparing the crop's aspect against each candidate. This is the step that buys the pixel density. The garment now occupies most of a native resolution canvas instead of a third of a downscaled frame, and the pleats have enough pixels to survive encoding into latents. **The tell:** count what fraction of the canvas the garment gets. If the crop hasn't moved it from a third of the frame to most of the canvas, you haven't bought any density, you've just moved the image around. --- ### Step 3 — Inpaint the crop and nothing else The inpainting model runs on the high fidelity crop. Same mask, same conditioning, same expert weights, applied to a much better input. ```python # pipeline/context_window.py (continued) INFERENCE = { "primary_steps": 30, "refinement_steps": 10, "sampler": "euler_ancestral", "scheduler": "beta", "vae_dtype": "fp16", "mask_dilation_px": 12, # bleeding zone, tuned per garment in 5..15 } def run_window( image: np.ndarray, mask: np.ndarray, pipe, conditioning, ) -> np.ndarray: """Inpaint the garment region at native resolution and return the crop.""" window = roi_from_mask(mask) target = pick_native(window) crop = resize(image[window.y0:window.y1, window.x0:window.x1], target) crop_mask = resize(mask[window.y0:window.y1, window.x0:window.x1], target) crop_mask = dilate(crop_mask, radius=INFERENCE["mask_dilation_px"]) generated = pipe( image=crop, mask_image=crop_mask, num_inference_steps=INFERENCE["primary_steps"], **conditioning, ).images[0] generated = pipe( image=generated, mask_image=crop_mask, num_inference_steps=INFERENCE["refinement_steps"], strength=0.25, **conditioning, ).images[0] print(f"inference at {target}, window {window.size}") return np.asarray(generated) ``` Thirty primary steps, then ten refinement steps at strength 0.25 over the same mask. The dilation is the same bleeding zone we use everywhere, a radius somewhere between 5 and 15 pixels depending on the garment. It lets the model see skin adjacent to the garment, so the boundary between fabric and body is something it generates rather than something we impose afterwards. For reference on cost, a full body try-on takes about 12 seconds on an RTX 4090. *That's a lab measurement on a single image, not a production figure, and I don't have a production figure to give you.* **The tell:** if you're cleaning up the fabric-to-skin boundary by hand after the fact, the dilation is too small. That boundary should be inside the region the model generated. --- ### Step 4 — Composite across the mask, not the window Downscale the generated crop back to the window's original size and paste it into the full resolution frame. The delivered asset keeps the resolution of the source photograph. Only the region inside the window has been touched. This is where the mechanism stops being straightforward, because pasting a rectangle of generated pixels into a photograph produces a visible rectangle. There are two distinct causes and they need different treatment. | seam cause | what you see | treatment | |---|---|---| | Hard edge at the window boundary | A faint rectangular outline a brand's art director finds in about four seconds | Feather the alpha over the dilated mask, not over the window rectangle, so the transition falls in the bleeding zone the model already blended | | Resampling mismatch | A change in grain across the boundary, the generated region slightly smoother than its surroundings | Match grain, or use the padding to push the boundary into flatter skin and backdrop | The second one is the harder of the two. The crop was generated at native resolution and then downscaled; the surrounding frame was never resampled. Those two regions have different noise characteristics and different high frequency content, and even with a perfect alpha the eye picks up the change. Downscaling is a low pass filter, so what comes back is smoother than the photograph it's landing in. **The tell:** if you can find a rectangle in the output, you feathered over the window instead of over the mask. --- ### What breaks it **The seam is caught by a human, late.** Right now a boundary artefact is found by a person looking at the image, which means it's found inconsistently and after the asset already exists. It should be a check that runs on every generated asset before anyone sees it. I haven't built that check, so treat the seam as a known unmeasured risk rather than a solved problem. **The mechanism raises a ceiling, it does not raise the floor.** Output scales to 4K for e-commerce print use, and quality was best with 2K input where the source garment photography was already high quality. That's worth saying plainly, because we spent a while confused by inconsistent results before recognising that the variable was input photography quality. A soft input garment produces a soft output garment at any resolution you like. **One window is an assumption, not a guarantee.** Everything above describes the single window case. When the garment doesn't fit in one crop, the window strategy has to generalise, and that generalisation has its own failure surface. --- ### What this post doesn't cover How many windows get used for a given garment, how they're placed, how much they overlap, and how the blend across multiple windows is handled. Those choices exist and they matter for the garments that don't fit in one crop. They're not something I'm writing up. Take the four steps as the description of the single window case, which is the case that covers most e-commerce shots. --- ### When the context window is the wrong choice The mechanism buys pixel density by refusing to spend compute on pixels nobody asked to change. If that description doesn't match your image, it buys you nothing and costs you a seam. - **The garment already fills the frame.** The saving comes from the backdrop, floor and headroom you're declining to process. Crop a close-up and the window is the frame, with a compositing boundary added for free. - **You have the VRAM.** A 48GB L40S is comfortable on the full frame. If your deployment target is a datacentre card and will stay one, running the frame whole avoids the entire seam problem. - **The structure you care about is short range.** Tiling breaks pleats because a pleat's correct value depends on content further away than the tile. Texture, print and colour don't have that property, and a mechanism built for long range continuity is overhead when nothing you're generating is continuous. - **Your reference photography is soft.** No inference resolution manufactures detail that was never in the reference. Fix the input first. - **A visible boundary is unacceptable and you can't place it well.** The second seam cause has no clean fix, only mitigation. If the window edge has to fall across high frequency detail, you're relying on grain matching rather than on geometry. --- ### The shift Cropping is a modelling decision, not a preprocessing detail. rho changes what the model knows about the body it's dressing, and it belongs in the same review as the sampler and the step count. I had it filed under image handling for far too long. The generation was never the hard part after we had the adapters. Getting generated pixels to sit inside a photograph without announcing themselves took longer than getting the pleats right. - **Treat the crop as a model parameter.** rho sits next to the sampler, not next to the file I/O. - **Spend your debugging time at boundaries.** That's where the failures that kill you live. - **Design against the smaller card.** The 24GB constraint made the pipeline run in more places, and it improved the architecture rather than compromising it. The concrete next step, if you take one thing from this: make the seam something a machine checks, not something a person notices. > A boundary artefact found by an art director is a boundary artefact you shipped. --- Source: https://himanshuat.com/blogs/fitting-a-saree-into-24gb ======================================================================== --- title: "Two streams: why one diffusion pass couldn't do virtual try-on" description: "Masking a torso and inpainting a garment demos well in an afternoon. Getting back the exact garment, with its weave and its print placement intact, is a different problem. One conditioning stream could not hold both the body and the reference, so Flux-VTON+ runs two." date: "March 18, 2025" url: "https://himanshuat.com/blogs/flux-fill-and-redux-two-streams" --- # Two streams: why one diffusion pass couldn't do virtual try-on Mask the torso, prompt for a red saree, run Flux Fill. You get a red saree, and it's usually a decent one. It is not the saree the brand shipped us. For a catalogue, that gap is the entire product. The outputs looked fine right up until someone on the brand side opened the file at full size. > Text is a bad channel for texture. You can write two hundred words describing a weave and the model will still pull it toward whatever weave dominated its training data. I spent a couple of weeks trying to fix that inside one conditioning stream before accepting it could not be fixed there. ### Chapter 0 — What a conditioning stream owns A conditioning stream is a path by which information reaches the sampler. Flux Fill takes an image, a mask, and a prompt, and fills the masked region with something that agrees with the surrounding pixels. That "something that agrees" is the whole problem. Given a torso, a pose, and the words "red saree", the model produces *a* red saree. Nothing obliges it to produce *the* red saree, because nothing in the conditioning path is carrying the zari border or the print repeat. Redux is a second path. It encodes the reference garment itself, so the specific weave has somewhere to enter. The division of labour is clean once you write it down: | Stream | Owns | What it cannot do | |---|---|---| | Flux Fill | *where* — body, pose, lighting on skin, the garment/not-garment boundary | Know which garment | | Flux Redux | *what* — weave, print, structure of the reference | Know where the shoulder is | Neither one can be asked to do the other's job. --- ### Step 1 — Name the failure mode before you name the fix Two things go wrong, and they are mirror images. **Style Drift** is what inpainting alone gives you. The output matches the reference in colour and rough silhouette and loses everything else. A knit sweater comes back as a flat red shirt. A hand-blocked print comes back as a generic floral. Colour survives because colour is low frequency and cheap to get right. Structure is where the money is, and structure is what goes first. **The Amputated Hand** is what structural guidance alone gives you. Condition strongly on the reference garment with a weak spatial anchor and you get a beautiful picture of the garment and a bad picture of a person wearing it. The drape ignores where the shoulder actually is. The hem lands at a plausible hem height rather than this model's hem height. Hands resting on a hip get repainted as fabric. It is the single fastest way to get an image rejected by a brand. **The tell:** look at your rejects and sort them into those two piles. If they land in one pile, you are missing one stream. If they land in both, you are missing both. --- ### Step 2 — Cut the mask with SAM2, then dilate it into a bleeding zone The mask is upstream of everything, and a bad mask poisons both streams at once. Threshold-based and parsing-based masks were not precise enough at the garment boundary, so we moved to SAM2 with a hierarchical prompt set and dilated the result. ```python # mask_garment.py """Produce the inpainting mask for a source image. SAM2 gives a tight garment boundary; morphological dilation turns that boundary into a "bleeding zone" the sampler is allowed to repaint, so the new fabric can meet skin without a visible seam. """ import cv2 import numpy as np import torch from sam2.build_sam import build_sam2 from sam2.sam2_image_predictor import SAM2ImagePredictor GARMENT_PROMPTS = ["shirt", "tshirt", "top", "saree", "dress"] # Bleeding zone radius in pixels, measured at the crop resolution the sampler # will actually see. 5 for tight-fitting tops, up to 15 for loose drape. DILATION_PX = 9 # Sanity floor: a mask covering less than this fraction of the frame usually # means SAM2 latched onto a sleeve or a shadow rather than the garment. MIN_MASK_FRACTION = 0.04 def build_predictor(checkpoint: str, config: str) -> SAM2ImagePredictor: """Load SAM2 once per worker; the encoder pass is the expensive part.""" model = build_sam2(config, checkpoint, device="cuda") return SAM2ImagePredictor(model) def segment_garment(predictor, image_bgr: np.ndarray, prompts=GARMENT_PROMPTS): """Run the hierarchical prompt set and keep the highest-scoring mask.""" image_rgb = cv2.cvtColor(image_bgr, cv2.COLOR_BGR2RGB) predictor.set_image(image_rgb) best_mask, best_score, best_prompt = None, -1.0, None for prompt in prompts: masks, scores, _ = predictor.predict(text=prompt, multimask_output=True) top = int(np.argmax(scores)) if scores[top] > best_score: best_mask, best_score, best_prompt = masks[top], float(scores[top]), prompt return best_mask.astype(np.uint8), best_score, best_prompt def dilate(mask: np.ndarray, radius_px: int = DILATION_PX) -> np.ndarray: """Grow the mask by radius_px with an elliptical structuring element.""" k = 2 * radius_px + 1 element = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (k, k)) return cv2.dilate(mask, element, iterations=1) def prepare_mask(predictor, image_bgr, radius_px=DILATION_PX): raw, score, prompt = segment_garment(predictor, image_bgr) coverage = float(raw.mean()) if coverage < MIN_MASK_FRACTION: raise ValueError(f"mask too small for '{prompt}': coverage={coverage:.3f}") final = dilate(raw, radius_px) grown = float(final.mean()) - coverage print(f"prompt={prompt} sam2_score={score:.3f} coverage={coverage:.3f} " f"radius={radius_px}px grown_by={grown:.3f}") return final ``` The dilation radius looks like a nuisance parameter and isn't. Set it too tight and the sampler never sees a skin pixel adjacent to the garment, so it has no evidence about the lighting on the body and paints an edge that reads as a cutout pasted on. Set it too wide and you erase the evidence you need: the shoulder line, the neck, the exact point where the arm leaves the torso. The model will happily invent a new shoulder, and an invented shoulder is worse than a hard seam because nobody notices it until the image is on a product page. Two things make the radius annoying to pin down. It is in pixels rather than a fraction of the body, so it has to be set at whatever resolution the crop is handed to the sampler, not at the resolution of the original file. And the right value depends on the garment. Something tight against the body wants a narrow zone because there is very little skin to blend into. Something with real drape wants a wide one, because the true boundary moves with the fabric and SAM2's tight boundary is only one plausible version of it. **5 to 15 pixels** covered everything we shot. *I never found a principled way to pick inside that range.* It stayed a per-category constant that somebody set by eye. **The tell:** zoom to the mask edge at full resolution. A visible seam means the zone is too tight. A shoulder that has moved means it is too wide. --- ### Step 3 — Encode the reference through Redux, not a plain CLIP encoder Redux is where the specific garment enters. The distinction I care about is against a plain CLIP vision encoder, which compresses the reference into something close to a caption in vector form. "A photo of a red shirt" is enough to steer colour and category and nowhere near enough to reproduce a print repeat. Redux keeps structural information through the bottleneck, which is why it can carry thread pattern and logo placement. ```python # reference_stream.py """Encode the reference garment into the structural conditioning vector. C_style = ReduxEncoder(SigCLIP(I_ref)) The garment is cut out before encoding. Background pixels in a reference shot are conditioning noise: a studio backdrop pushes the encoder toward "studio backdrop" and steals capacity from the weave. """ from dataclasses import dataclass import torch from PIL import Image @dataclass class ReferenceStreamConfig: redux_strength: float = 0.85 # scales C_style before cross-attention cutout_feather_px: int = 3 # soften the RMBG alpha edge square_pad: bool = True # pad rather than crop; aspect distortion # rotates print repeats and it shows class ReferenceStream: def __init__(self, siglip, redux, rmbg, config: ReferenceStreamConfig): self.siglip = siglip self.redux = redux self.rmbg = rmbg self.config = config def cutout(self, ref: Image.Image) -> Image.Image: """Isolate the garment on a transparent channel with RMBG.""" alpha = self.rmbg(ref) garment = ref.convert("RGBA") garment.putalpha(alpha.filter_edge(self.config.cutout_feather_px)) if self.config.square_pad: garment = pad_to_square(garment) return garment @torch.inference_mode() def encode(self, ref: Image.Image) -> torch.Tensor: """Return C_style, the conditioning injected into cross-attention.""" garment = self.cutout(ref) pixel_values = self.siglip.preprocess(garment).to("cuda", torch.bfloat16) # SigCLIP produces the visual features; Redux turns them into tokens # the Flux transformer can attend to alongside the text embedding. features = self.siglip(pixel_values) c_style = self.redux(features) c_style = c_style * self.config.redux_strength print(f"c_style tokens={tuple(c_style.shape)} " f"strength={self.config.redux_strength}") return c_style ``` The `redux_strength` scalar turned into the most-touched knob in the pipeline. Push it up and fidelity to the reference improves until it starts fighting the inpainting stream, at which point the garment stops respecting the pose and you see the reference's own drape stamped onto a body standing differently. Push it down and Style Drift comes back. There is a usable band, it is garment-dependent, and the failure on the high side is much easier to miss than the failure on the low side. A high-strength output looks crisp and detailed while being subtly wrong about the body. The RMBG cutout is not cosmetic. Reference shots arrive as flat lays, on hangers, or on a different model entirely, and whatever is behind the garment gets encoded along with it unless you remove it. **The tell:** if your reference encoder is a plain CLIP tower, describe the output you are getting. If the description fits any garment of that colour and category, the print repeat never made it through the bottleneck. --- ### Step 4 — Merge both streams into one sampler, then refine at low denoise Two conditioning paths, one sampler, and a second low-denoise pass over the same mask: ```mermaid flowchart LR R[reference] --> C[RMBG cutout] C --> S[SigCLIP + Redux] I[source image] --> M[SAM2 mask] M --> D[dilate] S --> F[Flux Fill sampler] D --> F F --> P{soft edge?} P -->|refine| F P -->|done| V[VAE decode] ``` ```python # run_vton.py """Single try-on pass: Fill conditioning on the masked body, Redux conditioning on the reference garment, one sampler, one refinement pass.""" SAMPLER = { "sampler_name": "euler_ancestral", "scheduler": "beta", "cfg": 3.5, } PRIMARY = {**SAMPLER, "steps": 30, "denoise": 1.0} REFINE = {**SAMPLER, "steps": 10, "denoise": 0.28} # The VAE runs in FP16 regardless of transformer precision. Lower precision # decoding produced colour blotching on large flat fabric areas, which is # the worst possible artefact for a solid-colour garment. VAE_DTYPE = "fp16" def try_on(pipe, source, mask, reference, ref_stream, prompt: str): """Run the two-stream pass and return the composited result.""" c_style = ref_stream.encode(reference) latents = pipe.fill( image=source, mask_image=mask, prompt=prompt, style_conditioning=c_style, **PRIMARY, ) # Second pass at low denoise. This is not a quality upscale; it exists to # settle the mask boundary and the high-frequency texture that the first # pass leaves slightly soft. latents = pipe.fill( image=latents, mask_image=mask, prompt=prompt, style_conditioning=c_style, **REFINE, ) image = pipe.decode(latents, vae_dtype=VAE_DTYPE) print(f"primary_steps={PRIMARY['steps']} refine_steps={REFINE['steps']} " f"sampler={SAMPLER['sampler_name']}/{SAMPLER['scheduler']}") return image ``` Euler Ancestral with the Beta scheduler was picked by comparison rather than by theory. The ancestral noise helps with fabric texture, which is high-frequency and slightly stochastic, and deterministic samplers tend to render it too smoothly. **Thirty steps for the primary pass and ten for refinement** is where quality stopped improving in a way I could see. The refinement pass at low denoise is doing boundary work more than anything else. Drop it and the seam at the mask edge starts showing up often enough to notice. **The tell:** turn the refinement pass off for a batch. If nothing changes at the mask edge, your primary pass is already settling the boundary and the second pass is just cost. --- ### Step 5 — Say out loud which parts are baselines The paper cites ControlNet and T2I-Adapter, and people assume they are in the stack. They aren't, and they never were. They are related work, and SDXL with a ControlNet was one of the two baselines we measured against. Spatial control over pose or depth is real, but it does not solve texture transfer, which is the actual VTON problem. Reference-only ControlNets do attempt texture, and in our testing they bled colour and degraded the style. The gap in the benchmark is mostly a global-garment gap. SSIM on Global Traditional garments came in at **0.45** for SDXL with ControlNet against **0.85** for the two-stream pipeline, while the Western Casual numbers were much closer together. Pose control was never the thing standing between us and a shippable saree. **The tell:** if half the questions you get about your architecture are people describing your related-work section back to you, the paper needs a sentence, not the reader. --- ### What breaks it **Sheer fabrics are still bad.** Lace and chiffon need the model to blend skin tone *with* fabric rather than paint fabric *over* skin, and the pipeline renders them close to opaque. I don't have a fix, only an awareness that we routed those SKUs away from the automated path. **Extreme poses degrade badly.** Standing and seated are well covered, anything acrobatic isn't, and the failure is a warp rather than an obvious break, so it doesn't trip any check. **The two knobs interact.** Redux strength and dilation radius are not independent. A wider bleeding zone gives the model more freedom, which makes a high Redux strength more likely to run away with the drape. When someone reported a regression, the cause was usually one of those two moving for a new garment category and nobody noticing the other one now needed to move too. --- ### When two streams is the wrong choice The second stream buys garment identity and costs you a reference encoder, a background-removal step, and a knob that has to be retuned per category. If you don't need identity, you are paying for all of that and getting nothing. - **You need a plausible garment, not a specific one.** Mood boards, concept shots, anything where "a red saree" is the requirement. Flux Fill alone does that in an afternoon, and Style Drift is not a defect when nobody is comparing against a reference. - **The garment is a flat solid colour.** Colour is low frequency and survives the text path already. The reference stream earns its cost on weave, print repeat, and logo placement. - **You cannot cut the reference out cleanly.** The stream carries whatever is behind the garment. A reference you can't isolate feeds the encoder studio backdrop instead of fabric, and that is worse than not having the stream. - **Your catalogue is mostly sheer fabric or extreme poses.** Both fail here regardless of how many streams you run. Route those away from the automated path rather than tuning against them. - **Nobody owns the per-category constants.** Dilation radius and Redux strength are hand-set, they interact, and they drift when a new category arrives. Without a person who notices, the pipeline degrades quietly. --- ### The shift Most of the early time went into prompt engineering that could never have worked, because the information I was trying to send does not survive the text encoder. The single-stream version wasn't undertuned. It was underdetermined, and no amount of sampler work was going to tell it which saree to produce. - **Match the channel to the shape of the information.** Text steers category and colour. Structure needs a visual path. - **Split responsibility instead of tuning one component harder.** Give each stream a job it can actually do and stop asking either to do both. - **Watch the boring parameters.** Dilation radius is a one-line morphological op that decides whether a brand accepts the image. The next thing to try is deriving the dilation radius from the segmentation rather than fixing it per category. SAM2 already gives a confidence signal at the boundary, and the places where it is least certain are exactly the places where the fabric moves. > That would remove the last hand-set number in the masking stage, which is also the number that broke production most often. --- Source: https://himanshuat.com/blogs/flux-fill-and-redux-two-streams ======================================================================== --- title: "From a ComfyUI graph to an API a brand's team can call" description: "Our try-on pipeline lived as a ComfyUI graph with implicit state, hand-picked seeds, and custom nodes tracking whatever was on main that week. Turning it into something another company's engineers could call meant writing down which code ran, which seed, when the work happens, and what we refuse." date: "September 09, 2025" url: "https://himanshuat.com/blogs/from-comfyui-graph-to-an-api" --- # From a ComfyUI graph to an API a brand's team can call In one week we got a 4K studio flatlay of a lehenga, a phone photo of a kurta on a hanger, and a PNG with a transparent background. All three came from people who reasonably assumed a try-on API takes a picture of clothing. The alpha channel surfaced as a Python traceback in a terminal I was watching, and as a timeout to the caller. Running a diffusion workflow yourself is easy to get away with. You keep the graph open in a browser tab, you nudge a denoise value when the fabric looks plastic, you re-run a node until the mask stops eating the model's hand, and nobody ever finds out how much of the final image came from you sitting there. > An API takes you out of the loop. Once the first beta keys went out to GettoIndia and EcomBuddha, the workflow had to behave the same way on a Tuesday afternoon as it did the night I tuned it. ### Chapter 0 — A wrapper exposes a tool, a contract states four promises The first version took about a day. ComfyUI already exposes an HTTP endpoint that accepts a workflow in its API JSON form, so I wrote a small Python service that loaded our saved graph, substituted the model image and the garment image into the two `LoadImage` nodes, posted it, and polled for the result. In a demo this is indistinguishable from a real product. As something another team depends on, it fell apart in four separate ways. | What the wrapper leaked | What the caller saw | |---|---| | Custom node versions floated. SAM2 masking, RMBG, and the LoRA merge were community nodes, installed the way everyone installs them, by cloning from git. | Nothing in the request told you which code had produced your image. | | The server held state we weren't tracking. A running ComfyUI process keeps models resident in VRAM, including a merged model built by `model_merge_lora` at some point in the past. | Edit the merge weights without restarting cleanly and a request can be served by weights that no longer match anything on disk. | | Failures came back as prose. A short-side resolution the graph didn't like, an alpha channel where we expected RGB, an image where SAM2 found no garment at all. | A traceback in my terminal. A timeout in theirs. | | Nothing stated what a caller was allowed to send. | The flatlay, the hanger photo, and the transparent PNG, all in the same week. | Resident VRAM state is fine when you're the only user and you know what you just changed. It is not fine when the person calling you is debugging their own storefront. The wrapper wasn't wrong so much as it was honest about the wrong thing. It exposed our tool. What an engineering team on the other side needs is a contract, and the contract turned out to be four separate promises: **which code ran, which random draw was used, when the work happens, and what we agree not to attempt.** --- ### Step 1 — Version the artifacts, not the graph The thing we versioned was never the workflow JSON on its own. It was the workflow plus every artifact the workflow reaches for, written down as a single file with content hashes. ```json // manifests/vton-2025-09-a.json { "manifest_id": "vton-2025-09-a", "created": "2025-09-02", "graph": { "file": "graphs/flux_vton_plus.api.json", "sha256": "9f2c1e0b7a4d3f88c5b6e2a1d0c9f4b3e8a7d6c5b4a39281706f5e4d3c2b1a09" }, "nodes": [ { "repo": "comfyui-sam2", "commit": "4b1e9c2" }, { "repo": "comfyui-rmbg", "commit": "a07d331" }, { "repo": "comfyui-essentials", "commit": "c19f4ab" } ], "models": [ { "role": "base", "file": "flux_fill.safetensors", "sha256": "1a2b3c4d" }, { "role": "redux", "file": "flux_redux.safetensors", "sha256": "5e6f7a8b" }, { "role": "vae", "file": "ae.safetensors", "sha256": "9c0d1e2f", "dtype": "fp16" }, { "role": "lora_drape", "file": "drape_physics_r32.safetensors", "sha256": "3a4b5c6d" }, { "role": "lora_occlusion", "file": "occlusion_depth_r32.safetensors", "sha256": "7e8f9a0b" } ], "merge": { "node": "model_merge_lora", "lambda_drape": 0.6, "lambda_occlusion": 0.4 }, "sampling": { "sampler": "euler_ancestral", "scheduler": "beta", "steps_primary": 30, "steps_refine": 10 }, "context_window": { "roi_padding": 0.2, "mask_dilation_px": 9, "inference_resolution": [1344, 1344] } } ``` Two things follow from writing it this way. The merge ratio stops being folklore that lives in my head and becomes part of the published surface. That matters because **0.6 for draping and 0.4 for occlusion was an empirical choice**, and someone would eventually ask why their output changed when we touched it. And the API version is no longer a number I increment when I feel like it. A request either names a manifest or gets the current default, and the response always echoes back the manifest it actually ran. When a caller reports that "the model got worse", the first question has an answer in the payload. We kept old manifests servable for as long as the weights fit on disk. That turned out to matter more than any other decision in this post. **The tell:** if a caller reports a regression and your first move is to check your own shell history, you shipped a graph, not a manifest. --- ### Step 2 — Derive the seed from the request ComfyUI gives every sampler node a seed and a control mode, and in interactive use the useful mode is randomize, because you're rolling for a good result. Through an API that behaviour is unacceptable. Two identical requests returning two different images means a caller cannot cache, cannot diff, and cannot write a regression test against you. ```python # api/seeding.py import hashlib import json from dataclasses import dataclass @dataclass class TryOnRequest: """A single try-on job as it arrives from a caller.""" request_id: str model_image_key: str garment_image_key: str category: str manifest_id: str seed: int | None = None def derive_seed(req: TryOnRequest) -> int: """Return the seed for this request. If the caller supplied a seed we use it untouched, so that replaying a request replays the draw. If they didn't, we derive one deterministically from the request contents rather than calling random(). Retrying the same job after a worker crash then lands on the same seed instead of quietly producing a different garment drape. """ if req.seed is not None: return req.seed % (2 ** 32) payload = json.dumps( { "model": req.model_image_key, "garment": req.garment_image_key, "category": req.category, "manifest": req.manifest_id, }, sort_keys=True, ).encode() return int.from_bytes(hashlib.sha256(payload).digest()[:4], "big") def apply_seed(graph: dict, seed: int) -> dict: """Write the seed into every sampler node and disable randomisation.""" for node in graph.values(): if node.get("class_type") in {"KSampler", "KSamplerAdvanced"}: node["inputs"]["seed"] = seed node["inputs"]["noise_seed"] = seed node["inputs"]["control_after_generate"] = "fixed" return graph if __name__ == "__main__": req = TryOnRequest( request_id="req_01J9", model_image_key="in/abf/model_204.png", garment_image_key="in/abf/saree_881.png", category="dresses_one_pieces", manifest_id="vton-2025-09-a", ) print("derived seed:", derive_seed(req)) print("stable across calls:", derive_seed(req) == derive_seed(req)) ``` The honest caveat goes in the docs next to this, not in a footnote. Same seed plus same manifest plus the same class of GPU gets you the same image. Move between GPU generations or driver versions and floating point accumulation order shifts, so you get an image that is visually the same and not bit identical. *I'd rather write that down than let a caller discover it while diffing PNGs.* The other half of seed handling is that the seed goes in the response. Brands wanted a specific generated look approved by their creative team and then reused across a set of SKUs. Without the seed in hand they were approving something they couldn't ask for again. **The tell:** if a caller can't replay yesterday's image, you're selling them a slot machine with a REST interface. --- ### Step 3 — Design for the unit of work you actually have A full-body try-on took roughly **12 seconds on an RTX 4090** with the model already resident. *That figure is a lab measurement of a single image on one machine, and I'm deliberately not turning it into a production latency claim, because we never ran the study that would let me.* What it does tell you is the shape of the system. The unit of work is seconds, and while it runs, a GPU is doing nothing else. That rules out request and response. The API accepts a job, returns a 202 with an id, and the caller either polls or gives us a webhook. ```mermaid flowchart LR A[submit] --> B{under limit} B -->|no| R[rejected] B -->|yes| Q[202 queued] Q --> W[worker lease] W -->|crash| Q W --> D[result] D --> P[poll or webhook] ``` ```python # api/queue.py import json import time import uuid import redis r = redis.Redis(decode_responses=True) QUEUE_KEY = "vton:pending" LEASE_TTL_SECONDS = 180 MAX_INFLIGHT_PER_KEY = 4 def submit(api_key: str, payload: dict) -> dict: """Admit a job or reject it, then return the caller-visible job record.""" inflight = r.scard(f"vton:inflight:{api_key}") if inflight >= MAX_INFLIGHT_PER_KEY: return {"status": "rejected", "reason": "concurrency_limit"} job_id = f"job_{uuid.uuid4().hex[:12]}" record = { "job_id": job_id, "api_key": api_key, "status": "queued", "manifest_id": payload["manifest_id"], "submitted_at": time.time(), "payload": payload, } r.hset(f"vton:job:{job_id}", mapping={"data": json.dumps(record)}) r.sadd(f"vton:inflight:{api_key}", job_id) r.lpush(QUEUE_KEY, job_id) return {"job_id": job_id, "status": "queued", "queue_depth": r.llen(QUEUE_KEY)} def claim(worker_id: str) -> dict | None: """Block until a job is available, then take a lease on it. The lease is what makes a worker crash survivable: if the worker dies mid-inference the lease expires and a sweeper puts the job back on the queue. Because the seed is derived from the request, the retry produces the same image the first attempt was going to produce. """ job_id = r.brpoplpush(QUEUE_KEY, f"vton:leased:{worker_id}", timeout=30) if job_id is None: return None r.setex(f"vton:lease:{job_id}", LEASE_TTL_SECONDS, worker_id) raw = r.hget(f"vton:job:{job_id}", "data") record = json.loads(raw) record["status"] = "running" r.hset(f"vton:job:{job_id}", mapping={"data": json.dumps(record)}) return record if __name__ == "__main__": print(submit("key_gettoindia", {"manifest_id": "vton-2025-09-a"})) print(submit("key_ecombuddha", {"manifest_id": "vton-2025-09-a"})) ``` Everything about the product side fell into place around that. Next.js on the front, NestJS holding the REST surface and the job records in Postgres, Redis for the queue and the leases, S3 for inputs, outputs and logs, and the Python services doing nothing but pulling work and running the graph. Making the inference layer a queue consumer rather than an HTTP server also meant it stopped mattering where it ran, which is a separate post. **The tell:** queue depth became the number we watched, and it was the only number worth putting on a dashboard, because every backend problem surfaces there first. If you can't name yours right now, you don't have a queue, you have a backlog. --- ### Step 4 — Write the refusals down The strategy was always specialized models per garment family rather than one generic try-on, with five categories planned: Dresses and One-Pieces, Tops, Bottoms, Intimates and Swimwear, and Co-ords and Sets. The API had to be honest about which of those actually had a model behind it at any given moment, and refuse the rest rather than produce something a brand would have to reject. ```yaml # api/surface.yaml categories: supported: - dresses_one_pieces - tops - bottoms planned: - intimates_swimwear - coords_sets input_requirements: model_image: color_space: rgb # alpha channels rejected, not flattened min_short_side: 1024 max_faces: 1 garment_image: background: any # RMBG runs before conditioning on_body_reference: preferred rejections: category_unsupported: http: 422 hint: "This category has no specialised model yet. See categories.planned." no_garment_found: http: 422 hint: "Segmentation returned an empty mask for the requested category." sheer_fabric_unsupported: http: 422 hint: "Lace and chiffon currently render opaque. Not accepted." pose_out_of_distribution: http: 422 hint: "Non-standing, non-seated poses are outside the trained range." ``` The last two rejections are the ones I'd argue for hardest. Sheer fabrics and unusual poses were known failure modes from our own evaluation, and an API that accepts them anyway is choosing to bill someone for an image its authors already know is wrong. > Refusing a request costs you a conversation once. Shipping a bad render costs you the account. **The tell:** if your evaluation named a failure mode and your endpoint still accepts it, the 422 you didn't write is a support ticket you already own. --- ### What breaks it **We hadn't defined idempotency.** A caller retrying a network timeout got a second job, a second GPU-second, and eventually a support email asking why one upload produced two images. Adding an optional idempotency key on submit fixed it in an afternoon and should have been there on day one. **Webhooks were worse than polling for a while.** Our first implementation retried on any non-200, and a partner's staging endpoint was returning 500 for a week. We were hammering them politely and constantly. **Pinning has no expiry.** One partner pinned an old manifest, upgraded nothing for a month, then filed what they described as a quality regression. It was our own older behaviour, reproduced exactly as promised. That was simultaneously the system working and a lesson that pinning without an expiry date is a way of preserving your worst output forever. --- ### When a contract is the wrong choice None of this pays for itself while the graph is still yours. - **You are the only caller.** Resident VRAM state and floating node versions are fine when you know what you just changed. The manifest is bookkeeping until somebody else is debugging their storefront against your output. - **You can't keep the old weights on disk.** We kept old manifests servable for as long as the weights fit. A pinned manifest you can't actually serve is a promise you've already broken. - **The caller needs bit-identical output across hardware.** Derived seeding gets you the same image on the same class of GPU. Across GPU generations or driver versions it gets you visually the same and not bit identical, and no amount of contract writing changes that. - **The unit of work isn't seconds of exclusive GPU.** The 202-plus-webhook machinery, the leases, the sweeper, the concurrency cap: all of it exists because one request owns a card for twelve seconds. If your work returns inside a normal request, you're buying operational surface for nothing. - **You haven't evaluated the model yet.** The refusal rules in `surface.yaml` came from failure modes we'd already found. Without an evaluation you don't know what to refuse, and a guessed 422 blocks work the model could have done. --- ### The shift The graph file was the least interesting thing in the manifest. The model hashes and the merge ratios were what actually determined the image, and those were the parts that had never been written down anywhere. Deterministic seeding is worth more than fast seeding. Deriving the seed from the request rather than a RNG made retries safe, made caching possible for the caller, and made bug reports reproducible, all from about fifteen lines of code. - **Version artifacts, not code.** Hash the weights, publish the merge ratios, echo the manifest back in every response. - **Derive the seed, don't draw it.** A request that can't be replayed can't be debugged by the person paying for it. - **Treat a refusal as a feature.** The failure modes you found during evaluation are the cheapest possible source of refusal rules. The next piece of this is per-category manifests. Once each garment family has its own merged model, one manifest per API version stops making sense, and the routing decision (which model, which LoRA pair, which resolution) has to move into the contract instead of living inside the worker. I'd rather make that change while we have two partners than after we have twenty. --- Source: https://himanshuat.com/blogs/from-comfyui-graph-to-an-api ======================================================================== --- title: "Part 1: Beyond Sequential Chains - Your First Agentic Workflow with LangGraph" description: "An LCEL chain that asks a clarifying question has nowhere to put the answer. It already ran to completion. LangGraph fixes that by making state explicit and letting the graph loop back on itself." date: "March 29, 2024" url: "https://himanshuat.com/blogs/getting-started-with-langgraph-agentic-workflows" --- # Part 1: Beyond Sequential Chains - Your First Agentic Workflow with LangGraph The agent decides your query is too vague, generates a clarifying question, and hands it back. Then the chain ends. There is nowhere for your answer to go, because the pipeline that asked the question has already run to completion and kept nothing. That is not a prompt problem. It is a shape problem. > Statelessness is not a missing feature you patch around. It is the property that decides whether you are writing a pipeline or an agent. I spend most of my time on AI research, mostly replicating cognitive architectures and Mixture-of-Experts (MoE) systems, and I keep running into the limits of current LLM frameworks at exactly this seam. ### Chapter 0 — What a graph adds to a chain LangChain Expression Language (LCEL) composes LLM calls into sequential pipelines. Define a prompt, pipe it to a model, pipe the output to a parser, pipe that to another model. ```python from langchain_core.prompts import ChatPromptTemplate from langchain_openai import ChatOpenAI from langchain_core.output_parsers import StrOutputParser # A simple LCEL chain prompt = ChatPromptTemplate.from_messages([ ("system", "You are a helpful assistant."), ("user", "{question}") ]) model = ChatOpenAI(temperature=0) output_parser = StrOutputParser() chain = prompt | model | output_parser # print(chain.invoke({"question": "What is the capital of France?"})) ``` This works for what it is. Single-shot tasks with a known sequence are exactly what it was built for. Hard problems don't run like that. You iterate, reflect, ask clarifying questions, branch, explore a couple of avenues, and converge. Throughout, you hold *state*: your current understanding, past attempts, the problem context. Every step shifts it. | | LCEL chain | LangGraph | |---|---|---| | State across steps | none, each invoke starts clean | explicit schema, updated per node | | Direction | one-directional | branches, loops, revisits | | Next step chosen by | the pipe order you wrote | a function of the current state | LangGraph gives you stateful, cyclical computation. It is an abstraction over finite state machines, which is a good fit for managing complex agent behaviour. I lean hard toward raw APIs and custom TypeScript over heavy Python frameworks, and I would still rather build this with minimal Python or Rust/Go services keeping state in Redis. But it is less kitchen-sink than vanilla LangChain and more focused on the graph execution engine, which is the right thing to focus on. The agent we're building takes a user query, decides whether it is clear enough to answer, and answers it if so. If not, it asks a clarifying question and loops back to the decision. --- ### Step 1 — Define the state before you write a single node The state schema is how information passes between nodes. A `TypedDict` gives you clarity and type-safety. ```python from typing import TypedDict, Annotated, List, Union import operator # Define the state schema class AgentState(TypedDict): """ Represents the state of our agent's workflow. This state is passed between nodes and updated. """ query: str clarifications: Annotated[List[str], operator.add] # Append new clarifications answer: Union[str, None] needs_clarification: bool # Flag to decide next step ``` `Annotated[List[str], operator.add]` is the part that earns its keep. It tells LangGraph that when a node returns a new list for `clarifications`, it should *append* to the existing list rather than overwrite it. Historical context survives the turn. Fields without that annotation get replaced. `query`, `answer`, and `needs_clarification` are all last-write-wins, which is what you want for them. **The tell:** if a field gets clobbered when you wanted it accumulated, you left the reducer off. The schema is where that decision lives, not the node. --- ### Step 2 — Make each node one function with one job A node is a Python function that takes the current `AgentState` and returns a partial state to merge. Nothing more exotic than that. ```python import os from langchain_core.prompts import ChatPromptTemplate from langchain_openai import ChatOpenAI # Ensure your OPENAI_API_KEY is set as an environment variable # os.environ["OPENAI_API_KEY"] = "YOUR_API_KEY" llm = ChatOpenAI(model="gpt-4o-mini", temperature=0.2) # Using a faster, cheaper model ``` The first node is a router. It asks the model whether the query is clear enough, and returns nothing but the flag. ```python def decide_action(state: AgentState) -> AgentState: """ Decides whether the query needs clarification or can be answered directly. """ print("---DECIDE ACTION NODE---") query = state["query"] # Simple prompt to let the LLM decide decision_prompt = ChatPromptTemplate.from_messages([ ("system", "You are an intelligent router. Based on the user's query, decide if it is clear enough " "to provide a direct answer, or if it requires more clarification. " "Respond 'CLARIFY' if more information is needed, otherwise respond 'ANSWER'.\n\n" "Current query: {query}\n" "Existing clarifications (if any): {clarifications}" ), ("user", "Decision for: {query}") ]) decision_chain = decision_prompt | llm | StrOutputParser() raw_decision = decision_chain.invoke({ "query": query, "clarifications": "\n".join(state.get("clarifications", [])) }) decision = raw_decision.strip().upper() print(f"Decision: {decision}") return {"needs_clarification": (decision == "CLARIFY")} ``` Note what it does *not* do. It doesn't answer, it doesn't clarify, and it doesn't name the next node. It writes one boolean and stops. The second node answers, reading any clarifications already accumulated in state. ```python def answer_query(state: AgentState) -> AgentState: """ Answers the user's query directly. """ print("---ANSWER QUERY NODE---") query = state["query"] answer_prompt = ChatPromptTemplate.from_messages([ ("system", "You are a helpful assistant. Provide a concise and accurate answer to the following query. " "If clarifications were previously provided, use them to inform your answer.\n\n" "Query: {query}\n" "Clarifications: {clarifications}" ), ("user", "{query}") ]) answer_chain = answer_prompt | llm | StrOutputParser() response = answer_chain.invoke({ "query": query, "clarifications": "\n".join(state.get("clarifications", [])) }) print(f"Answer: {response}") return {"answer": response} ``` The third generates the clarifying question. It returns a single-element list, which the `operator.add` reducer appends, and it clears any stale answer. ```python def ask_clarification(state: AgentState) -> AgentState: """ Asks the user for clarification based on the current query. """ print("---ASK CLARIFICATION NODE---") query = state["query"] clarification_prompt = ChatPromptTemplate.from_messages([ ("system", "You are a polite assistant. Formulate a concise and clear question to get more information " "from the user regarding their query. Focus on what's missing.\n\n" "Original Query: {query}\n" "Previous Clarifications (if any): {clarifications}" ), ("user", "What clarification is needed for: {query}") ]) clarification_chain = clarification_prompt | llm | StrOutputParser() clarifying_question = clarification_chain.invoke({ "query": query, "clarifications": "\n".join(state.get("clarifications", [])) }) print(f"Clarification needed: {clarifying_question}") return {"clarifications": [clarifying_question], "answer": None} # Clear previous answer if any ``` Three functions, three responsibilities, each testable on a plain dict without a graph anywhere near it. Any competitive programmer recognises the pattern: break a hard problem into composable units. **The tell:** if a node function needs to know which node runs next, you have put the decision in the wrong place. That belongs on an edge. --- ### Step 3 — Put the decision on the edge, not inside the node `StateGraph` wires the nodes together. This is where the non-linear behaviour actually comes from. ```python from langgraph.graph import StateGraph, END # Create the graph instance workflow = StateGraph(AgentState) # Add the nodes workflow.add_node("decide_action", decide_action) workflow.add_node("answer_query", answer_query) workflow.add_node("ask_clarification", ask_clarification) # Set the entry point (where the graph starts) workflow.set_entry_point("decide_action") # Define edges: how nodes connect # If 'decide_action' determines no clarification is needed, go to 'answer_query' # Otherwise, go to 'ask_clarification' workflow.add_conditional_edges( "decide_action", # Source node lambda state: "clarify" if state["needs_clarification"] else "answer", # A function that determines the next node { "clarify": "ask_clarification", "answer": "answer_query" } ) # After answering, the workflow is complete workflow.add_edge("answer_query", END) # After asking for clarification, we loop back to 'decide_action' # The user provides new input, and the graph effectively restarts from 'decide_action' # with the updated state. For demonstration, we'll manually feed updated state. # For a real agent, this would typically involve waiting for user input. workflow.add_edge("ask_clarification", "decide_action") # Loop back! This is the key. # Compile the graph app = workflow.compile() ``` `add_conditional_edges` is the engine of the dynamic behaviour. The routing function is a pure function of state, so the branch is inspectable without running a model. The last edge is the one that makes this a graph rather than a chain drawn sideways. ```mermaid flowchart LR S[query] --> D{decide} D -->|answer| A[answer] D -->|clarify| C[clarify] C --> D A --> E[END] ``` **The tell:** if you can trace every path from the entry point to `END` without revisiting a node, you have written a chain with extra ceremony. --- ### Step 4 — Run it as a stream and watch the state move `stream` yields state updates as they happen, keyed by the node that produced them. For a workflow that may run several turns, that is what you want to watch rather than a single final blob. ```python # Initial query initial_state = { "query": "I need help with my Python code.", "clarifications": [], "answer": None, "needs_clarification": False # Initialized, will be updated by node } print("\n--- First Execution ---") # Invoke the graph with the initial state # The `stream` method is efficient for long-running agents, yielding state updates for s in app.stream(initial_state): print(s) print("---") # Let's assume the last state yielded is the final one final_state_after_first_run = list(app.stream(initial_state))[-1] # If clarification was asked, manually update the query for the next turn if final_state_after_first_run.get("ask_clarification") or final_state_after_first_run.get("decide_action", {}).get("needs_clarification"): print("\n--- Agent requested clarification. User provides more info. ---") # Simulate user updating the query based on clarification updated_query = "My Python code is a web scraper using BeautifulSoup and I'm getting a 'NoneType' error when trying to find an element." # Create new state for the next turn, merging updated query and existing clarifications # In a real app, this would be a new user message. Here we manually update `query` # LangGraph state update implicitly merges dictionaries. next_state = { **initial_state, # Preserve original structure "query": updated_query, "clarifications": final_state_after_first_run.get("ask_clarification", {}).get("clarifications", []) # Carry over the clarifying question asked } print("\n--- Second Execution with updated query ---") for s in app.stream(next_state): print(s) print("---") # Get the final answer final_state = list(app.stream(next_state))[-1] print(f"\nFinal Answer: {final_state.get('answer_query', {}).get('answer')}") else: final_state = final_state_after_first_run print(f"\nFinal Answer: {final_state.get('answer_query', {}).get('answer')}") ``` The user turn is simulated here. `updated_query` stands in for the message a real user would send back, and the clarifying question is carried forward into the next invocation by hand. Once the graph has more than a handful of nodes, look at it rather than reason about it. ```python from IPython.display import Image, display # Requires pygraphviz # If you don't have graphviz, skip this or install it: # pip install pygraphviz # sudo apt-get install graphviz (on Debian/Ubuntu) try: display(Image(app.get_graph().draw_png())) except Exception as e: print(f"Could not draw graph: {e}. Make sure pygraphviz and graphviz are installed.") ``` The rendering path needs system packages, not just a pip install: `pip install pygraphviz` plus `sudo apt-get install graphviz` on Debian or Ubuntu. The `try` block is there because that install fails often enough to be worth handling. **The tell:** if a node's name never appears as a key in the streamed updates, that node never ran, whatever the final state looks like. --- ### What breaks it **The loop back has no human in it.** `add_edge("ask_clarification", "decide_action")` returns control straight to the router with `query` unchanged. The clarifying question gets appended to state, and then the same router re-judges the same query. Nothing in the compiled graph pauses for an answer. That's why the demo above feeds updated state by hand rather than letting the cycle run, and it is the single biggest gap between this and a real agent. **The router returns a bare string and the comparison is exact.** `decision == "CLARIFY"` after `strip().upper()`. Any other output, including a polite sentence wrapped around the word, falls through to the answer branch. The failure is silent: you get an answer to a question the agent should have asked about. **The driver runs the graph twice.** `list(app.stream(initial_state))[-1]` re-invokes the whole workflow after the `for` loop has already streamed it once. In a demo that costs a few extra model calls. In anything with side effects, it is a duplicate execution you didn't ask for. --- ### When LangGraph is the wrong choice The graph machinery is not free. It buys branching and state, and if you don't need those it is structure without a payoff. - **The task is single-shot with a known sequence.** Prompt, model, parser, done. LCEL is clean and direct for exactly this, and adding a state schema and three node registrations to it buys nothing. - **Every branch function is deterministic.** If no node calls a model and all your routing is `lambda state: ...` over plain fields, you have written a state machine with a framework attached. Ordinary Python control flow does that, and you can read it in one screen. - **Overhead is the thing you're optimising.** LangGraph is still an abstraction over the calls underneath it. For minimum overhead, which is always my goal, a custom graph executor on raw API calls with explicit state storage (Redis for inter-service communication) and compile-to-native routing logic gives you fine control over resource allocation and latency. That control matters most in high-throughput research. - **You need the loop to stop and wait.** As written, the cycle in this graph never yields to a user. If your workflow genuinely blocks on external input, you need something outside the compiled graph to hold state between turns before the cycle is worth building at all. --- ### The shift State is the whole difference. Without an explicit schema there is no dynamic, iterative behaviour, only a pipeline that runs and forgets. The `AgentState` `TypedDict` and the `operator.add` reducer are what make evolving context survive a node boundary. The graph is what gives an agent something like agency. Moving from a linear chain to a cyclical graph lets it decide, branch, and come back, which is closer to how a hard problem actually gets solved. That is the foundation for MoE architectures, where expert routing and iterative refinement are the entire point. - **Write the state schema first.** The reducers on it decide what your nodes can even remember. - **Keep nodes ignorant of routing.** One function, one job, one partial state returned. - **Check for a cycle.** No cycle means you built a chain and paid for a graph. This is Part 1. Next: harder decision points, and folding external tools into these loops. --- Source: https://himanshuat.com/blogs/getting-started-with-langgraph-agentic-workflows ======================================================================== --- title: "Graph agents and when the graph earns its keep" description: "A graph framework hands you a second control flow language layered on top of Python, and you pay for it in state design, checkpoint size, and debugging. I built the same agent both ways under constraints I couldn't negotiate. The line where the graph wins turned out to be narrow, and most of my systems fell on the other side of it." date: "May 12, 2026" url: "https://himanshuat.com/blogs/graph-agents-when-the-graph-earns-its-keep" --- # Graph agents and when the graph earns its keep One of my nodes had a static `add_edge` leaving it and also returned `Command(goto=...)`. Both fired. The downstream node executed twice, which meant a duplicated paid tool call inside a run somebody was waiting on the output of. The afternoon that followed went into discovering that the framework had done exactly what it documents. > A graph framework is a second control flow language layered on top of Python. It earns that only where cycles, durability, and mid-run human approval arrive together. Take away any one of the three and the plain loop wins. Two constraints made the original decision for me, and neither was mine to argue with. Runs were long and expensive enough that a failure halfway through could not mean starting over from the first tool call. And a human had to be able to approve an action mid-run before it went any further. The delivery date was fixed too, which ruled out finding out the hard way and rebuilding later. So I built the same agent twice, once as a plain Python while loop and once on LangGraph. The loop was about ninety lines shorter. Only one of the two survived being killed mid-run. Understanding why took most of April, because it came down to a single architectural fact I had skimmed straight past in the docs. ### Chapter 0 — What a super-step is Almost every diagram of a LangGraph agent looks like a flowchart, so you assume the runtime is a DAG executor. Topologically sort the nodes, run them in order, done. That model is wrong, and every confusing thing about the framework follows from it being wrong. The execution engine is Pregel, the same message-passing model Google published for large-scale graph processing. Execution proceeds in **super-steps**. A node activates when it receives a message, runs, writes its state updates, and votes to halt. The run terminates when all nodes are inactive and no messages are in transit. Nodes that fan out in a single super-step run in parallel with each other. Nodes that run in sequence span separate super-steps. Hold onto that and a pile of confusing behaviour stops being confusing. Parallel writes to the same state key conflict because they happen inside one super-step, not because of a race in your code. A whole node re-runs on resume, rather than resuming mid-function, because the super-step is the atomic unit of progress. And `recursion_limit` counts super-steps, so it has almost nothing to do with how many LLM calls you made — a single super-step can contain four parallel nodes, each making three model calls. Once I stopped reading the graph as a flowchart and started reading it as a scheduler, I could predict its behaviour instead of discovering it. --- ### Step 1 — Design the state schema before you write a node You define state as a TypedDict, and the annotations on that dict do far more work than the node functions do. ```python # graph.py import operator from typing import Annotated, TypedDict from langchain_core.messages import AnyMessage from langgraph.graph import END, START, StateGraph from langgraph.graph.message import add_messages class ResearchState(TypedDict): """State carried across every super-step of the research graph. Every key in here is checkpointed on each super-step, so anything that does not need to survive a resume should not live here. """ messages: Annotated[list[AnyMessage], add_messages] findings: Annotated[list[str], operator.add] question: str verdict: str class ResearchInput(TypedDict): """What a caller is allowed to pass in.""" question: str class ResearchOutput(TypedDict): """What a caller gets back, which is far less than the full state.""" verdict: str findings: list[str] def plan(state: ResearchState) -> dict: """Turn the question into a search plan, recorded as a message.""" ... def search(state: ResearchState) -> dict: """Run one retrieval pass and append whatever it turned up.""" ... def critique(state: ResearchState) -> dict: """Decide whether the findings so far actually answer the question.""" ... def route(state: ResearchState) -> str: """Edge function. Loop back into search, or stop.""" return END if state["verdict"] == "sufficient" else "search" builder = StateGraph( ResearchState, input_schema=ResearchInput, output_schema=ResearchOutput, ) builder.add_node("plan", plan) builder.add_node("search", search) builder.add_node("critique", critique) builder.add_edge(START, "plan") builder.add_edge("plan", "search") builder.add_edge("search", "critique") builder.add_conditional_edges("critique", route, ["search", END]) graph = builder.compile() ``` The topology that buys is a cycle with a gate on it, which is the shape a while loop gives you for free and a DAG executor cannot express at all. ```mermaid flowchart LR ST([start]) --> P[plan] P --> S[search] S --> C{critique} C -->|insufficient| S C -->|sufficient| E([end]) ``` Two things in that file are load-bearing. `input_schema` and `output_schema` let the graph carry a fat internal state while exposing a narrow contract, which matters more than it sounds like it does once other code starts calling your graph. And `.compile()` is mandatory: the builder is not runnable, and compilation is where the channel wiring and validation happen. The annotations are reducers. `add_messages` merges message lists by id, so re-emitting a message updates it rather than duplicating it. `operator.add` concatenates. A key with no reducer gets last-write-wins semantics, which is fine right up until two nodes write it in the same super-step. **The tell:** if you cannot say which keys a node needs to read *after* a resume, you have not designed a state schema. You have collected local variables in the wrong scope. --- ### Step 2 — Say your reducer's conflict policy out loud Here is the failure that taught me the most. Fan two nodes out of `START` and have both write the same unreduced key. ```python # fanout_conflict.py from typing import TypedDict from langgraph.graph import START, StateGraph class State(TypedDict): """Deliberately reducer-free, which is the entire point of the example.""" topic: str notes: list[str] def west(state: State) -> dict: """One of two nodes activated in the same super-step.""" return {"notes": ["west says p99 looks fine"]} def east(state: State) -> dict: """The other one. Same key, same super-step, different opinion.""" return {"notes": ["east says p99 does not look fine"]} builder = StateGraph(State) builder.add_node("west", west) builder.add_node("east", east) builder.add_edge(START, "west") builder.add_edge(START, "east") graph = builder.compile() graph.invoke({"topic": "p99", "notes": []}) # langgraph.errors.InvalidUpdateError: At key 'notes': Can receive only one # value per step. Use an Annotated key to handle multiple values. ``` The docs file that under the error code `INVALID_CONCURRENT_GRAPH_UPDATE`, and the fix everyone reaches for is to slap `Annotated[list[str], operator.add]` on the key. It makes the error go away. It does not make the problem go away. **Adding a reducer chooses a policy.** Before the reducer, two concurrent writers to one key were a hard error, loudly, at the exact moment of conflict. After the reducer, they are an order-dependent result. `operator.add` concatenates in whatever order the writes land within the super-step, so the contents of `notes` now depend on which node finished first, and nothing in the type signature says so. You converted a crash into a silent ordering dependency and called it a fix. That generalises. Picking `operator.add` says "order does not matter to me". Picking a custom reducer that sorts by a stable key, or that merges by node id, says something stronger and usually more accurate. I now write custom reducers for anything two nodes can touch, even trivial ones, purely so the policy is written down somewhere. **The tell:** if you added a reducer to make an error message stop and you cannot state its conflict policy in one sentence, you did not resolve the race. You hid it. --- ### Step 3 — Reach for `Send` when the fan-out width is dynamic Static edges cannot express a width you only learn at runtime. `Send` can. Each `Send` carries its own private payload to the target node, so the workers do not all read the same shared blob. ```python # fanout_send.py import operator from typing import Annotated, TypedDict from langgraph.graph import END, START, StateGraph from langgraph.types import Send class MapState(TypedDict): """Parent state. `summaries` is the only key workers write.""" documents: list[str] summaries: Annotated[list[str], operator.add] report: str class WorkerState(TypedDict): """Private per-Send payload. Never merged back except via the reducer.""" document: str def dispatch(state: MapState) -> list[Send]: """Fan out one worker per document inside a single super-step.""" return [Send("summarise", {"document": doc}) for doc in state["documents"]] def summarise(state: WorkerState) -> dict: """Summarise exactly one document. Runs in parallel with its siblings.""" return {"summaries": [f"summary of {state['document'][:40]}"]} def reduce_report(state: MapState) -> dict: """Collapse every worker summary into one report.""" return {"report": "\n".join(state["summaries"])} builder = StateGraph(MapState) builder.add_node("summarise", summarise) builder.add_node("reduce_report", reduce_report) builder.add_conditional_edges(START, dispatch, ["summarise"]) builder.add_edge("summarise", "reduce_report") builder.add_edge("reduce_report", END) graph = builder.compile() ``` Every one of those `summarise` invocations lands in one super-step, and every write to `summaries` goes through `operator.add`. ```mermaid flowchart LR D[dispatch] --> S1[summarise] D --> S2[summarise] D --> S3[summarise] S1 --> R[reduce report] S2 --> R S3 --> R ``` That is the exact situation from Step 2. The difference is that here I genuinely do not care about ordering, because the reduce step joins them anyway. Same mechanism, chosen deliberately instead of inherited. **The tell:** if you know the number of workers at build time, static edges are enough and `Send` is ceremony. `Send` is for the width you cannot write down until the run is already going. --- ### Step 4 — Keep the graph drawable when nodes route themselves `Command` lets a node update state and pick its own next hop in one return value. The interesting part is the type hint. ```python # supervisor.py from typing import Literal, TypedDict from langgraph.graph import END, START, StateGraph from langgraph.types import Command class TeamState(TypedDict): """Shared state for a supervisor and two workers.""" task: str draft: str review_notes: str def supervisor(state: TeamState) -> Command[Literal["writer", "reviewer", "__end__"]]: """Route the task and record the routing decision in one step. The Literal in the return annotation is what lets the graph be rendered statically. Drop it and the visualiser sees a dead end. """ if not state["draft"]: return Command(update={"task": state["task"].strip()}, goto="writer") if not state["review_notes"]: return Command(goto="reviewer") return Command(goto=END) def writer(state: TeamState) -> Command[Literal["supervisor"]]: """Produce a draft, then hand control back up.""" return Command(update={"draft": "..."}, goto="supervisor") def reviewer(state: TeamState) -> Command[Literal["supervisor"]]: """Review the draft, then hand control back up.""" return Command(update={"review_notes": "..."}, goto="supervisor") builder = StateGraph(TeamState) builder.add_node("supervisor", supervisor) builder.add_node("writer", writer) builder.add_node("reviewer", reviewer) builder.add_edge(START, "supervisor") graph = builder.compile() ``` `Command` also takes `graph=Command.PARENT` to hand control back out of a subgraph, and `resume=` to feed a value into an interrupted run. That `resume=` path is the human-in-the-loop story. It's the one capability that genuinely has no clean equivalent in a while loop, and it is most of why the graph won the build it won. The trap is the one from the top of this post. A node with a static `add_edge` leaving it that *also* returns `Command(goto=...)` fires both, and the downstream node executes twice. The docs warn about it. I walked into it anyway and paid for the tool call twice. **The tell:** render the compiled graph before you run it. If a routing node draws as a dead end, its `Literal` annotation is missing, and you are about to debug control flow you cannot see. --- ### Step 5 — Set durability and a step budget on day one Persistence has three modes, typed as `Durability = Literal['sync', 'async', 'exit']`. Sync writes the checkpoint before the next step starts. Async writes it while the next step is already executing. Exit persists only when the run exits. *The docs I read do not state which one is the default, so I won't either.* Set it explicitly and you never have to care. ```python # run.py from langgraph.checkpoint.sqlite import SqliteSaver from langgraph.errors import GraphRecursionError from graph import builder def run(question: str, thread_id: str) -> dict: """Run one research thread with sync durability and a tight step budget. recursion_limit counts super-steps, not model calls. The default of 1000 is a runaway-protection number, not a budget. """ with SqliteSaver.from_conn_string("checkpoints.db") as saver: graph = builder.compile(checkpointer=saver) config = { "configurable": {"thread_id": thread_id}, "recursion_limit": 24, } try: return graph.invoke({"question": question}, config, durability="sync") except GraphRecursionError: snapshot = graph.get_state(config) return { "verdict": "budget_exhausted", "findings": snapshot.values.get("findings", []), } ``` The default `recursion_limit` is 1000 super-steps, and blowing through it raises `GraphRecursionError`. Treat that as a smoke alarm rather than a budget. A two-node ping-pong between a planner and a critic will spend real money long before it reaches 1000, and it will do so cheerfully, because nothing in the graph knows the difference between progress and oscillation. The graceful alternative is the managed `RemainingSteps` channel, which lets a node see how much budget is left and wrap up deliberately instead of dying at the limit. The costs I did not anticipate are mostly about state. Everything in state is checkpointed every super-step, so **checkpoint size scales with how much you're carrying, not with how many decisions you've made.** `add_messages` grows monotonically by design, so a long agent thread means every super-step writes an ever-larger blob. Two hundred super-steps over a fat state is a lot of writes for very little information, and the write amplification is invisible until you look at the size of the checkpoint table. What fixed it was being ruthless about what deserves to be in state at all. Retrieved document text moved to a content store, with state holding only the ids. Intermediate scratch values that no downstream node reads got deleted rather than left lying around because they were convenient during development. **The tell:** if a key in state is never read by a node that runs after a resume, you are paying to checkpoint a local variable on every super-step. --- ### Step 6 — Debug by diffing state history, not by hunting for a stack trace Debugging here is different in a way nobody warns you about. There is no stack trace that spans the run, because the run is not a call stack. It's a sequence of super-steps, and by the time you notice something is wrong the frames you'd want to inspect were discarded several steps ago. `get_state_history()` is the actual debugger, and reading it is a skill you have to deliberately acquire rather than something you pick up. The workflow that replaced breakpoints for me is dull and effective. Reproduce with a fixed `thread_id`. Walk the state history backwards until the state stops looking right. Note the super-step where a key first held a value nobody expected. That's your failing node, and it is frequently not the node that raised. **The tell:** if your first instinct on a bad run is to set a breakpoint, you are still debugging a call stack that does not exist. --- ### Alternatives that occupy different ground Three alternatives are worth knowing precisely, because they sit at different points rather than competing on the same one. | alternative | where control flow lives | what you trade | |---|---|---| | OpenAI Agents SDK | No graph. Handoff is a tool call named `transfer_to_`, picked like any other tool, bounded by `max_turns` raising `MaxTurnsExceeded` | Much less machinery, much less say over the path | | `pydantic-graph` | Nodes are dataclasses subclassing `BaseNode[StateT, DepsT, RunEndT]`; the edge set is the return type of `run()`, terminating at `End(value)` | Control flow checked by the type checker instead of by string node names | | Temporal-style durable execution | Deterministic workflow code replayed from an event history; activities execute once and have their results recorded | Cheap to store, and real determinism constraints on the code you write | Pydantic's own docs say graphs are "not the right tool for every job", which I appreciated. The line worth drawing across the whole table is this. A checkpointer stores state snapshots; deterministic replay stores the decision log. Snapshots are easy to reason about and expensive when state is large. Replay is cheap to store and imposes real constraints on your code. Neither is strictly better, and knowing which one you need is more useful than knowing either framework's API. --- ### What breaks it **The state schema is a database migration nobody treats as one.** Renaming a state key silently loses the persisted value on existing threads. The old key is still in the checkpoint, the new key reads as absent, nothing errors, and the graph carries on with a default. I found that out on threads that were already in flight, which is the worst way to find it out. **Two ways of leaving a node both fire.** A static `add_edge` plus a `Command(goto=...)` from the same node runs the downstream node twice. Everything in that node happens twice, including whatever it pays for. **The step limit cannot tell progress from oscillation.** `recursion_limit` counts super-steps and defaults to 1000. A planner and a critic bouncing off each other will burn a real bill before they get anywhere near that number, and the graph will report nothing wrong the entire time. --- ### When a graph is the wrong choice The graph earned its keep in exactly one of my two builds, and only because that build needed cycles, durability, and human approval mid-run at the same time. - **You need only cycles.** That's a while loop. You already know how to write one, and it will be shorter. - **You need only durability.** That's a job queue with a status table. The checkpointer is a lot of machinery to buy for a column. - **You need only human-in-the-loop.** That's a state machine with a pause in it. - **Your control flow is fixed and you know the order.** A model call at each step of a known sequence is a workflow, not an agent. Anthropic's guidance says the blunter version: start with the simplest thing and add agentic machinery only when the simpler approach falls short. Most production systems are workflows wearing agent branding. Mine were. - **Your state is large and your decisions are few.** Snapshotting a fat state on every super-step is the expensive way to buy durability. Deterministic replay stores the decision log instead, and that's a different axis, not a worse framework. - **You want the model to choose the path.** Then the graph's routing is machinery you're paying for and not using, and a handoff-as-tool-call SDK is the honest shape. It's the combination that gets ugly to hand-roll. Pausing a cyclic process at an arbitrary point and resuming it later on a different machine is the thing the checkpointer and the super-step boundary actually solve. --- ### The shift I stopped asking whether a graph would be nice to have and started asking which of three properties I was buying. That question is answerable in an afternoon. The one I was asking before took me most of April. - **Read the runtime as a scheduler.** Super-steps explain parallel write conflicts, whole-node re-runs, and what `recursion_limit` counts. - **Treat every reducer as a concurrency policy.** If you cannot state the policy, you hid a race rather than fixing it. - **Put a key in state only if a post-resume node reads it.** Everything else is a local variable you're paying to checkpoint. If you're on the fence right now, do this before you pick. Write down the longest path your system takes and count how many steps on it are decisions the model makes, versus steps whose order you already know. > If the answer is mostly the latter, write the loop. If you need to interrupt that path, get a human to approve something, and come back an hour later on a different process, compile the graph and set `durability="sync"` on day one, because retrofitting persistence onto a live schema is the migration you don't want. --- Source: https://himanshuat.com/blogs/graph-agents-when-the-graph-earns-its-keep ======================================================================== --- title: "LangChain Agents 101: Building Your First Autonomous Tool-User From Scratch" description: "An agent is a prompt, a regex, a dict lookup, and a for loop. This builds one out of `requests` and `re` so that all four are yours to read, and names the places each of them breaks." date: "April 05, 2024" url: "https://himanshuat.com/blogs/langchain-agents-101-building-first-autonomous-tool-user" --- # LangChain Agents 101: Building Your First Autonomous Tool-User From Scratch The model ends its turn with `Action: search` on the last line, no trailing newline. Your parser wants `Action: (\w+)\n`. It matches nothing. The agent records "LLM did not provide a valid Action", burns an iteration, and asks again. Nothing in that path throws. Nothing in your logs points at a model. That failure is not in the model and it is not in the tool. It is in the two characters of contract between them, which is exactly the layer a framework's `initialize_agent` puts behind a wall. > The framework isn't hiding complexity from you. It's hiding the four places your agent actually breaks: the prompt, the parse, the dispatch, and the exit. You probably landed here searching for "LangChain Agents 101." This post does not use LangChain. Frameworks like it lower the barrier to entry and then become bloatware, hiding the details you need for performance tuning and real understanding. My goal isn't to string APIs together, it's to understand and eventually replicate intelligence, and obscured abstractions get in the way of that. Read the title as "Agentic Systems 101, from first principles." ### Chapter 0 — What ReAct actually is ReAct is a prompt-engineering technique. It is not a LangChain invention and there is no class you need to inherit from to use it. It names four things the model and your code trade back and forth: * **Thought:** the model's internal monologue, planning the next step. * **Action:** the tool the model decides to use. * **Action Input:** the parameters for that tool. * **Observation:** the result the tool returns, which feeds back into the model for the next `Thought`. That is the whole cognitive loop: observe, think, act. The LLM is the reasoning engine, the frontal lobe. The tools are the sensory-motor system, plain Python callables that do external work like searching the web, querying a database, or running code. In a Mixture-of-Experts setup those tools stand in for specialized expert networks or external peripherals, which is why I care about controlling exactly how they interact, what data they see, and how fast they process it. Everything below is four sections of one file. Nothing in it imports a framework. --- ### Step 1 — Call the model over HTTP and own the response shape Start with the smallest thing that can be called an LLM interface: a `requests.post` and a dictionary. Temperature goes to zero, because a stochastic planner on top of a string-matching parser is two sources of variance where you can afford one. ```python # agent.py import json import re import requests # For making actual API calls import os # To get API keys # --- 1. LLM Interface --- # Replace with your actual OpenAI API key or a compatible service OPENAI_API_KEY = os.getenv("OPENAI_API_KEY", "YOUR_MOCK_API_KEY") OPENAI_ENDPOINT = "https://api.openai.com/v1/chat/completions" # Or your local LLM endpoint def call_llm(prompt: str, model: str = "gpt-4o-mini") -> str: """Makes a raw API call to an OpenAI-compatible LLM.""" headers = { "Content-Type": "application/json", "Authorization": f"Bearer {OPENAI_API_KEY}", } payload = { "model": model, "messages": [{"role": "user", "content": prompt}], "temperature": 0.0, # Keep it deterministic for agentic behavior } try: # In a real scenario, handle async with 'httpx' or 'aiohttp' for performance response = requests.post(OPENAI_ENDPOINT, headers=headers, json=payload) response.raise_for_status() # Raise an exception for bad status codes result = response.json() return result['choices'][0]['message']['content'] except requests.exceptions.RequestException as e: print(f"Error calling LLM: {e}") return f"ERROR: Could not communicate with LLM: {e}" except KeyError: print(f"Error parsing LLM response: {response.json()}") return "ERROR: Malformed LLM response." # --- MOCK LLM CALL (for local testing without an actual API key) --- def mock_call_llm(prompt: str) -> str: """A mock LLM for local testing. Simulates ReAct responses.""" print("\n--- MOCK LLM CALL ---") if "what is the capital of France" in prompt.lower() and "Observation:" not in prompt: return "Thought: The user is asking for the capital of France. I should use the search tool.\nAction: search\nAction Input: capital of France" elif "capital of france" in prompt.lower() and "Observation:" in prompt: return "Thought: The search results clearly state that Paris is the capital. I have enough information now.\nFinal Answer: The capital of France is Paris." elif "who won the world series in 2020" in prompt.lower() and "Observation:" not in prompt: return "Thought: I need to find the winner of the 2020 World Series. The search tool is appropriate for this.\nAction: search\nAction Input: 2020 World Series winner" elif "2020 world series winner" in prompt.lower() and "Observation:" in prompt: return "Thought: The search results indicate the Los Angeles Dodgers won. I can now provide the final answer.\nFinal Answer: The Los Angeles Dodgers won the World Series in 2020." elif "current date" in prompt.lower(): return "Thought: The user is asking for the current date. I should use the get_current_date tool.\nAction: get_current_date\nAction Input: today" elif "current date" in prompt.lower() and "Observation:" in prompt: return "Thought: I have the current date from the tool. I can provide the final answer.\nFinal Answer: The current date is October 26, 2023." # This will be dynamic with real tool else: return "Thought: I am unable to answer this question with the available tools or based on prior context. I need to be more explicit or use a different tool. Final Answer: I cannot answer this question based on the provided information and tools." # Choose which LLM function to use LLM_FUNC = call_llm if OPENAI_API_KEY != "YOUR_MOCK_API_KEY" else mock_call_llm ``` Two things earn their place here. The `KeyError` branch exists because a malformed response body and a failed request are different problems and you want them distinguishable at the point of failure. The mock exists because a ReAct loop has a control-flow bug rate independent of the model, and you should be able to exercise the control flow without paying per token. **The tell:** if you cannot run your agent end to end with the network unplugged, every loop bug you find will cost an API call to reproduce. --- ### Step 2 — Make tools plain functions and let the docstring be the spec A tool is a callable and a description. There is no wrapper type, no registry, no decorator. ```python # --- 2. Tool Definitions --- # These are simple functions. No special LangChain Tool wrappers needed. def search_tool(query: str) -> str: """A search engine. Use this to answer questions about current events or facts. Input is the search query (e.g., "latest news"). """ print(f"DEBUG: Executing search for '{query}'...") # In a real system, you'd integrate with a search API (e.g., SerpApi, Google Custom Search). # For this example, we'll return a deterministic mock result. if "capital of France" in query: return "Observation: Paris is the capital and most populous city of France." elif "2020 World Series winner" in query: return "Observation: The Los Angeles Dodgers won the 2020 World Series, defeating the Tampa Bay Rays." else: return f"Observation: No specific search result found for '{query}'. (This is a mock search)." from datetime import datetime def get_current_date(input_str: str) -> str: """Returns the current date. Input is ignored (can be 'today' or empty).""" print(f"DEBUG: Executing get_current_date with input '{input_str}'...") return f"Observation: The current date is {datetime.now().strftime('%B %d, %Y')}." # Store tools in a dictionary for easy lookup and description generation TOOLS = { "search": { "func": search_tool, "description": search_tool.__doc__.strip() }, "get_current_date": { "func": get_current_date, "description": get_current_date.__doc__.strip() } } ``` The detail worth copying is `search_tool.__doc__.strip()`. The text the model reads to decide whether to call this function is the same text a developer reads above the implementation. They cannot drift apart, because they are one string. Note also that both tools take a single `str` and return a single `str`. That is not laziness, it is the type the prompt protocol can actually express. `Action Input` is one line of text. **The tell:** if your tool description lives in a config file, a decorator argument, or a JSON schema separate from the function body, you now have two specifications and only one of them is executed. --- ### Step 3 — Write the prompt that carries the contract The prompt is the interface definition. It states the format, enumerates the legal tool names from the same dict the dispatcher will use, and replays the history. ```python # --- 3. Agent Prompt Template --- # This is crucial for guiding the LLM to follow the ReAct pattern. def generate_agent_prompt(task: str, thought_history: list[str], current_observation: str, tools_dict: dict) -> str: """Generates the prompt for the LLM based on the task, history, and available tools.""" tool_descriptions = "\n".join([ f"{name}: {data['description']}" for name, data in tools_dict.items() ]) # Format the history to be clean for the LLM history_str = "\n".join(thought_history) prompt = f"""You are an AI assistant designed to answer questions and solve problems using tools. You have access to the following tools: {tool_descriptions} The current task is: "{task}" You should strictly follow this format: Thought: You must always think about what to do next. Action: The name of the tool to use, must be one of [{', '.join(tools_dict.keys())}] Action Input: The input string to the tool Observation: The result from the tool ... (This Thought/Action/Action Input/Observation cycle can repeat N times) Thought: I have enough information to provide a final answer. Final Answer: The ultimate answer to the original question. If you believe you have already answered the question, state the Final Answer. Current conversation history: {history_str} {f"Observation: {current_observation}" if current_observation else ""} Thought: """ return prompt ``` `', '.join(tools_dict.keys())` is the line that matters. The list of legal actions in the prompt is generated from the dict the loop looks up, so adding a tool cannot leave the two out of sync. Every other coupling in this file is a string you have to keep true by hand. The prompt ends on a bare `Thought:` so the model's first emitted token continues the format rather than deciding on one. **The tell:** if you cannot paste the exact string your model received into an editor and read it top to bottom, you do not know what you asked. Uncomment the prompt print and look at it once per feature. --- ### Step 4 — Own the loop and name every way out of it This is the part frameworks call an executor. It is a `for` loop with four regexes and a dict lookup. ```python # --- 4. The Agent Loop --- # This is where the magic happens – we manually control the ReAct cycle. def run_agent(task: str, max_iterations: int = 5) -> str: """Runs the ReAct agent loop.""" thought_history = [] current_observation = "" for i in range(max_iterations): print(f"\n--- Iteration {i+1}/{max_iterations} ---") prompt = generate_agent_prompt(task, thought_history, current_observation, TOOLS) # print(f"LLM Input:\n```\n{prompt}\n```") # Uncomment to see the full prompt llm_response = LLM_FUNC(prompt) print(f"LLM Output:\n```\n{llm_response}\n```") thought_history.append(llm_response) # Keep full LLM response for history # Parse LLM response for Thought, Action, Action Input thought_match = re.search(r"Thought: (.*?)(?=\nAction:|\nFinal Answer:|$)", llm_response, re.DOTALL) action_match = re.search(r"Action: (\w+)\n", llm_response) action_input_match = re.search(r"Action Input: (.*)", llm_response) final_answer_match = re.search(r"Final Answer: (.*)", llm_response, re.DOTALL) thought = thought_match.group(1).strip() if thought_match else "" action = action_match.group(1).strip() if action_match else "" action_input = action_input_match.group(1).strip() if action_input_match else "" if final_answer_match: print(f"Agent providing Final Answer!") return final_answer_match.group(1).strip() if action and action_input: if action in TOOLS: print(f"Agent executing tool '{action}' with input '{action_input}'...") tool_func = TOOLS[action]["func"] current_observation = tool_func(action_input) print(f"Tool Observation: {current_observation}") else: current_observation = f"Observation: Error: Unknown tool '{action}'. Available tools: {', '.join(TOOLS.keys())}" print(current_observation) else: current_observation = "Observation: Error: LLM did not provide a valid Action/Action Input, or reached an ambiguous state." print(current_observation) return f"Agent reached maximum iterations ({max_iterations}) without a Final Answer. Last thought: {thought_history[-1] if thought_history else 'None'}" ``` The branch structure is the whole design, and it is small enough to hold in your head: ```mermaid flowchart TD P[build prompt] --> L[call model] L --> C{final answer?} C -->|yes| R[return answer] C -->|no| A{known tool?} A -->|yes| T[run tool] A -->|no| E[error observation] T --> P E --> P ``` Five conditions, five outcomes, and each one costs something: | condition | what the loop does | cost | |---|---|---| | `Final Answer:` matched | returns immediately, ignores any action | ends the run | | action parsed, name in `TOOLS` | calls the function, stores its return | one iteration | | action parsed, name unknown | feeds an error string back as the observation | one iteration | | no action, no final answer | feeds "ambiguous state" back as the observation | one iteration | | `max_iterations` exhausted | returns a failure string with the last raw response | the whole budget | The three middle rows are the ones people miss. A recovery attempt and a successful tool call are charged at the same rate, so a model that misspells a tool name three times has spent the same budget as one that did three real steps. `max_iterations=3` is not three steps of work. Note also that `final_answer_match` is checked before the action branch, so a response containing both a `Final Answer` and an `Action` will answer and never call the tool. **The tell:** count the `return` statements and the observation-assignment branches in your loop. If that count is larger than the number of exit conditions you can name from memory, the loop is deciding things you have not decided. --- ### Step 5 — Run it against a task it should fail The last invocation here is the one worth keeping. Three tasks the tools cover, and one they do not. ```python # --- Example Usage --- if __name__ == "__main__": print("\n--- Running Agent for: 'What is the capital of France?' ---") result1 = run_agent("What is the capital of France?", max_iterations=3) print(f"\nFinal Result 1: {result1}") print("\n--- Running Agent for: 'Who won the World Series in 2020?' ---") result2 = run_agent("Who won the World Series in 2020?", max_iterations=3) print(f"\nFinal Result 2: {result2}") print("\n--- Running Agent for: 'What is the current date?' ---") result3 = run_agent("What is the current date?", max_iterations=3) print(f"\nFinal Result 3: {result3}") print("\n--- Running Agent for: 'Tell me a joke.' (Should fail gracefully) ---") result4 = run_agent("Tell me a joke.", max_iterations=2) print(f"\nFinal Result 4: {result4}") ``` "Tell me a joke" has no tool behind it. The interesting question is not whether the agent answers, it's which exit path it takes to give up: a `Final Answer` that declines, or the `max_iterations` string. Those are different bugs wearing the same output. **The tell:** if every task in your test file is one your tools can serve, you have tested the happy path and none of the four exits. --- ### What breaks it **The parser is the real interface, and it is stricter than the prompt says.** `Action: (\w+)\n` requires a newline after the tool name, so an action on the final line never matches. `\w+` quietly rejects any tool name with a hyphen or a dot. `Action Input: (.*)` stops at the end of the line, so a multi-line input is silently truncated to its first line. The prompt promises a format; the regexes enforce a narrower one, and the gap between the two is where the model looks obedient and the code disagrees. **Observations do not accumulate.** `thought_history` collects model responses, but `current_observation` is a single variable overwritten every iteration and only the latest one is rendered into the prompt. On turn three the tool result from turn one is gone. For a two-step task this never surfaces. For anything that has to hold two facts at once, the model is being asked to reason over evidence it can no longer see. **The `Observation:` prefix is written twice.** Each tool returns a string that already starts with `Observation: `, and `generate_agent_prompt` prepends `Observation: ` again when rendering it. The model reads `Observation: Observation: Paris is the capital...`. It usually copes, which is the problem: the loop is teaching the model a format the loop itself doesn't follow, and you'll find it by reading the assembled prompt rather than by watching anything fail. All three are one class of bug. Nothing here throws, so nothing here shows up as an error. The only instrument that finds them is the printed prompt. --- ### When writing the loop yourself is the wrong choice Owning all four layers is the point of this exercise, and it is not free. There are cases where it is the wrong trade. - **You need the integrations, not the loop.** The loop is forty lines and you can read it in a sitting. The connectors, retries, and vendor quirks a framework ships are the part that would take you months, and none of it is what you learn here. - **The task is a single tool call.** No branch, no second observation. The ReAct scaffold, the parse layer, and the iteration budget are pure overhead around what is a function call with a model in front of it. - **Your model already emits a machine-readable action.** Everything in Step 4's parse block exists because the contract is prose. If you can get the action and its arguments back as structured data, hand-written regexes over free text are a downgrade you are choosing. - **Somebody else has to maintain it.** A bespoke parser with an undocumented strictness gap is a fine thing to own alone and an expensive thing to hand over. The framework is sometimes the cheaper social choice rather than the cheaper technical one. For production I'd move this to TypeScript for type safety and performance, but the control-flow argument doesn't change with the language. --- ### The shift The five things this file teaches are not about agents. They're about which layer you are allowed to debug. ReAct is a prompt pattern, so the quality of your agent's planning is a property of a string you wrote. Tools are functions, so the quality of tool selection is a property of a docstring you wrote. When you own the loop, you can inspect every prompt, every response, every tool call and observation, which is what a long reasoning chain costs to diagnose. And the dependency list is `requests` and `re`. That granular control is not an aesthetic preference about bloat. Modular, specialized intelligence needs it: with expert modules in an MoE setup, you have to control exactly how they interact, what data they see, and how fast they process it. - **Read the assembled prompt before you tune anything.** Most agent bugs are visible in it and invisible in the traceback. - **Treat the parser as the contract, not the prompt.** The prompt is what you asked for; the regex is what you accept. - **Name every exit before you raise `max_iterations`.** Retries and real work cost the same iteration. Build this once yourself and the mechanics transfer to any agentic system, whatever the model and whatever the tools. It's the difference between driving a car and building an engine. --- Source: https://himanshuat.com/blogs/langchain-agents-101-building-first-autonomous-tool-user ======================================================================== --- title: "LLM as judge, what the numbers actually say" description: "I let a judge gate whether changes shipped, because the blog posts said it agreed with humans about 80% of the time. Six weeks in, I reported an improvement that was the judge flipping on rerun. This is what that 80% actually measures, and how little of it survives chance correction." date: "June 23, 2026" url: "https://himanshuat.com/blogs/llm-as-judge-what-the-numbers-actually-say" --- # LLM as judge, what the numbers actually say The judge decided whether a change shipped. I'd adopted it because the blog posts said it agreed with humans about 80% of the time. Six weeks later I reported an improvement that turned out to be the judge flipping on rerun. The number had already been passed on to someone who took it at face value, and correcting that in writing is a specific kind of unpleasant. > A judge that agrees with your humans 85% of the time is not 85% right. It is 85% similar to a group that only agrees with itself four times in five. That sent me back to the primary sources. The headline figure is doing far less work than the way it gets quoted suggests. ### Chapter 0 — What "agreement" is a measurement of The number comes from MT-Bench and Chatbot Arena, Zheng et al., NeurIPS 2023 Datasets and Benchmarks track, arXiv 2306.05685. Know the shape of the data before you quote the result. MT-Bench is 80 multi-turn questions across 8 categories with 3,000 expert votes from 58 expert labelers. Chatbot Arena contributed 30,000 conversations judged by 2,114 unique crowdsourced judges. Excluding ties, GPT-4 pairwise agreed with human experts 85% of the time, and single-answer grading on the second turn agreed 84%. Those are the numbers everyone repeats. The number almost nobody repeats sits right next to them in the same paper: **human-human agreement was 81 to 82%.** That comparison is the whole point, and getting it backwards is the most common error in judge writing. GPT-4 matches human-to-human agreement. It does not match ground truth, because on these tasks there is no ground truth to match. Two qualified humans disagree about one answer in five. The judge is not approximating a correct answer. It is approximating a distribution of human opinion, and the ceiling on that is your annotation guideline. If your guideline is vague, your humans will disagree, and a judge that perfectly reproduced your humans would still look unreliable. I found this clarifying and slightly deflating. A large chunk of "the judge is unreliable" turns out to be "my rubric is underspecified". --- ### Step 1 — Chance-correct before you believe any agreement number The most differentiating material I found is a June 2026 preprint, "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models", arXiv 2606.19544. Flag for anyone about to cite it: this is a v1 preprint and has not been peer reviewed. Treat the findings as strong signal, not settled fact. The scale is what makes it worth reading. 21 judges from 9 providers, evaluated across 3 benchmarks and 3 protocols, 118 runs, roughly 541,000 individual judgments. The headline finding, and the reason I rebuilt my evaluation: **kappa deflation is universal.** Between exact-match agreement and Cohen's kappa on MT-Bench, the gap runs 33 to 41 percentage points. Every claim of the form "our judge agrees with humans 85% of the time" overstates reliability by 30 to 40 points once you correct for chance agreement. Raw agreement counts the cases where the judge and the human both said "A wins" because A wins most of the time in your dataset. Kappa does not. Cohen's kappa is the minimum bar. The Landis and Koch (1977) bands are the usual reading, and they are conventions rather than laws — Landis and Koch themselves offered them as arbitrary but useful. | kappa | Landis and Koch reading | |---|---| | 0.21 to 0.40 | fair | | 0.41 to 0.60 | moderate | | 0.61 to 0.80 | substantial | | 0.81 to 1.00 | almost perfect | Krippendorff's alpha handles more than two raters and missing data, with the conventional cutoffs of 0.800 for reliable and 0.667 for tentative conclusions only. If you take one thing from this post, take the deflation gap. The 85% is real, and the reliability it implies is not. **The tell:** if the only agreement number you can quote for your judge is a raw percentage, you have not measured reliability yet. You have measured how skewed your label distribution is. --- ### Step 2 — Report per criterion, never an aggregate Reference points make the bands concrete. MT-Bench human-human agreement of 81 to 82% raw is the realistic ceiling for a subjective task. MAST's human annotators reached kappa 0.88 on a failure taxonomy, which is what a tight guideline looks like. A clinical LLM-jury study reported kappa 0.74 against human consensus, while per-criterion kappa ranged from 0.34 to 0.97. That last range is the reason I report per criterion and have stopped reporting an aggregate at all. A headline 0.74 hides a criterion sitting at 0.34, which in Landis and Koch terms is fair agreement, meaning barely better than the coin. Averaging is how you ship a judge that is excellent at four things and useless at the fifth without ever finding out which is which. ```python # agreement.py from collections import Counter def cohens_kappa(a: list[str], b: list[str]) -> float: """Chance-corrected agreement between two raters over the same items. The gap between this and raw agreement is the whole point. On skewed label distributions raw agreement can look excellent while kappa sits near zero, because both raters are mostly just guessing the majority. """ if len(a) != len(b) or not a: raise ValueError("rater vectors must be non-empty and equal length") n = len(a) observed = sum(1 for x, y in zip(a, b) if x == y) / n ca, cb = Counter(a), Counter(b) expected = sum((ca[k] / n) * (cb[k] / n) for k in set(a) | set(b)) if expected == 1.0: return 1.0 return (observed - expected) / (1 - expected) def per_criterion_report(labels: dict[str, tuple[list[str], list[str]]]) -> dict: """Kappa for every criterion separately, plus raw agreement alongside. Never collapse this to a mean. The point of the table is to expose the criterion where the judge falls apart, and a mean is designed to hide it. """ out = {} for criterion, (human, judge) in labels.items(): raw = sum(1 for x, y in zip(human, judge) if x == y) / len(human) kappa = cohens_kappa(human, judge) out[criterion] = { "raw_agreement": round(raw, 3), "kappa": round(kappa, 3), "deflation_points": round((raw - kappa) * 100, 1), } return out ``` That `deflation_points` field is there because I wanted to see the 33 to 41 point gap on my own data rather than take it on faith. I saw it. **The tell:** if your evaluation produces one number per run instead of one number per criterion, you cannot answer the question "which criterion is the judge bad at". Nobody can answer it for you either. --- ### Step 3 — Run every pair in both presentation orders Position bias is measured as consistency after swapping the order of the two answers. The 2023 numbers, default prompt against a prompt that only renamed the assistants: | judge | default prompt | renamed prompt | |---|---|---| | GPT-4 | 65.0% | 66.2% | | GPT-3.5 | 46.2% | 51.2% | | Claude-v1 | 23.8% | 56.2% | Read the last row again. Claude-v1's consistency more than doubled from merely renaming the assistants in the prompt. Nothing about the answers changed, nothing about the task changed, and the judge became a different judge. That is the sharpest prompt-sensitivity result in the paper and the reason I now treat any judge prompt as a component with its own version number. ```python # position_bias.py from dataclasses import dataclass from itertools import product @dataclass class PairJudgment: """One judgment plus the order the answers were shown in.""" task_id: str order: str # "ab" or "ba" winner: str # "a", "b", or "tie" def consistency(judgments: list[PairJudgment]) -> float: """Fraction of tasks where swapping the order did not change the winner. This is the metric MT-Bench reports. It is not accuracy: a judge can be perfectly consistent and perfectly wrong. Consistency below roughly 0.7 means the judge cannot resolve small differences at all, because the presentation order is a larger signal than the quality difference. """ by_task: dict[str, dict[str, str]] = {} for j in judgments: by_task.setdefault(j.task_id, {})[j.order] = j.winner complete = [v for v in by_task.values() if "ab" in v and "ba" in v] if not complete: raise ValueError("no task was judged in both orders") agree = sum(1 for v in complete if v["ab"] == v["ba"]) return agree / len(complete) def build_swapped_runs(tasks: list[str], answers: dict[str, tuple[str, str]]): """Emit every task twice, once in each presentation order. Running only one order is the single most common evaluation bug I see. It hides position bias completely, because there is nothing to compare. """ for task_id, order in product(tasks, ("ab", "ba")): a, b = answers[task_id] yield task_id, order, (a, b) if order == "ab" else (b, a) ``` **The tell:** if you have never judged the same pair twice in opposite orders, your position bias is not low. It is unmeasured. --- ### Which biases are still live Verbosity bias was tested in the same paper with a repetitive-list attack, where an answer is made longer without adding information, across 23 answers. | judge | fell for the repetitive-list attack | |---|---| | Claude-v1 | 91.3% | | GPT-3.5 | 91.3% | | GPT-4 | 8.7% | The spread is enormous, and it says the bias is a property of specific models rather than of the judging paradigm. The 2026 preprint carries the contrarian follow-up: verbosity bias measured below 0.011 across the cohort under a single pairwise rubric. Against 2023, where two of three models fell for the attack more than 90% of the time, that suggests verbosity bias has largely been trained out while position bias has not. The standard bias listicle that treats them as equally live concerns is out of date on one of them. Self-enhancement bias is the third. GPT-4 favoured its own outputs with roughly a 10% higher win rate, Claude-v1 by roughly 25%. The paper is careful to say this is suggestive rather than causally isolated, since the models also differ in ways that could explain part of it, and I'll be equally careful here. --- ### Step 4 — Give the judge a reference answer Table 4 of the same paper is the most practically useful thing in it. On math questions, across 20 judgments: | judge prompt | failure rate | |---|---| | default | 70% | | + chain-of-thought | 30% | | + reference-guided grading | 15% | Seventy to thirty to fifteen. That progression is the cleanest published argument for reference-guided grading, and it's why I now consider a judge without a reference answer to be a fallback rather than a design. Generating a reference is work, but it is work you do once per task, not once per evaluation run. **The tell:** if your judge is grading a math or code answer with nothing to compare it against, you are paying for an opinion when a comparison was available. --- ### Step 5 — Keep a canary set of content-free answers Separate from the bias literature, "One Token to Fool LLM-as-a-Judge" (arXiv 2507.08794) demonstrates that content-free "master keys" elicit false-positive rewards. The examples are as bare as they sound: a lone colon, a full stop, the phrase "Thought process:", and "Let's solve this problem step by step." No answer, no reasoning, no content. Rewards anyway, across GPT-o1, Claude-4, and multiple model scales. The paper's mitigation is Master-RM, a reward model trained with truncated outputs as adversarial negatives so that a plausible-looking prefix with nothing behind it gets scored as the failure it is. *The abstract gives no per-model false-positive rates, so I'm not going to put a figure on how often any specific model falls for it.* What I do now is keep a small canary set of content-free answers in every judge evaluation. If the judge passes any of them, the judge is broken, and nothing else I measure that day matters. ```python # canaries.py MASTER_KEYS = [ ":", ".", "Thought process:", "Let's solve this problem step by step.", ] def canary_report(judge, tasks: list[dict]) -> dict: """Score content-free answers and count how many the judge passes. Any pass here is disqualifying. A judge that rewards an empty prefix is not measuring answer quality, it is measuring surface plausibility, and every other number it produces inherits that. """ failures = [] for task in tasks: for key in MASTER_KEYS: verdict = judge(question=task["question"], answer=key) if verdict.get("pass"): failures.append({"task": task["id"], "key": key}) total = len(tasks) * len(MASTER_KEYS) return { "trials": total, "false_positives": len(failures), "examples": failures[:10], "usable": len(failures) == 0, } ``` **The tell:** if a lone colon scores a pass, every other number in your report is describing surface plausibility, not quality. --- ### Step 6 — Price a panel before you buy one big judge Replacing a single large judge with a panel of smaller ones is the alternative with the best cost story. PoLL, "Replacing Judges with Juries" (arXiv 2404.18796), uses a panel drawn from **disjoint model families**, which is the load-bearing detail: models from the same family share biases, so a panel of three siblings mostly amplifies one opinion. The panel is reported as more than 7 times less expensive than a single large judge and less prone to intra-model bias. Size is not the whole story either. Prometheus, a 13B fine-tuned evaluator, reports Pearson correlation with human evaluators across 45 rubrics: | evaluator | Pearson with humans | |---|---| | Prometheus (13B) | 0.897 | | GPT-4 | 0.882 | | ChatGPT | 0.392 | G-Eval, with a GPT-4 backbone, reports Spearman 0.514 with humans on summarization, and its own authors flag that it shows a preference for LLM-generated text. And JUDGE-BENCH, spanning 20 datasets and 11 LLMs, lands on the conclusion that works as the counterweight to every 80% headline: variance across models and datasets is substantial, and judges "should be carefully validated against human judgments before being used as evaluators". **The tell:** if your panel members share a model family, you have one judge with a higher bill. --- ### The pipeline I ended up with The judge is still in the pipeline and it still gates changes. What changed is what I let it decide on its own. It runs after the deterministic checks, with a reference answer, on binary criteria, in both presentation orders, with a canary set attached, and I report kappa per criterion rather than an aggregate agreement percentage. ```mermaid flowchart TD A[candidate output] --> B{deterministic check} B -->|answers it| C[free verdict] B -->|cannot| D[judge, order ab] B -->|cannot| E[judge, order ba] D --> F{same winner} E --> F F -->|yes| G[record verdict] F -->|no| H[unresolved] ``` The `unresolved` branch is the part I would have skipped a year ago. A flipped verdict is data about the instrument, and folding it into a majority vote throws that data away. --- ### What breaks it **Rerun stability is not quality, and it looks exactly like quality.** The 2026 preprint names this the consistency-bias paradox: production judges with test-retest reliability above 0.95 simultaneously show position bias above 0.10. The most reproducible judges are among the least valid ones. A judge that returns the same answer every time is reproducing its own biases faithfully, which is what reproducibility measures and all it measures. I had been using rerun stability as my proxy for quality, and signing off on results with it. **Leaderboards do not transfer.** Judge rankings shift by up to 14 positions across benchmarks in the same preprint. A judge ranked near the top on one benchmark can land near the bottom on another. Picking a judge off a leaderboard is therefore unsound, and I had done exactly that. **Binary criteria buy decisions and cost you resolution.** My practitioner preference is binary pass/fail over Likert scales, and the sourcing here is weaker than everything above, so treat it as opinion. Likert without worked exemplars collapses toward central scores, and judges do not share a latent notion of what separates a 3 from a 4. Two judges scoring the same answer 3 and 4 are not disagreeing about quality, they are disagreeing about the scale. The honest tension: G-Eval's probability-weighted normalization exists precisely because coarse integer scores are tie-heavy and give you no ranking resolution. If you need to rank twenty systems, binary labels will produce a wall of ties. Pick based on what you'll do with the output. --- ### When a judge is the wrong choice The clearest signal is a deterministic check existing. Anthropic orders verification as rules-based first and describes the judge as something that "isn't considered very robust". If a linter, a schema validator, or an exact-match test can answer the question, that answer is free and correct, and a judge is a worse version of it. Beyond that, five situations where I'd hold off. - **Before you've validated against human labels on your own data.** Published agreement numbers transfer poorly, and the 14-position ranking shift across benchmarks says so directly. - **Reference-free evaluation of output from the same model family.** The self-enhancement numbers are roughly 10% for GPT-4 and 25% for Claude-v1 on their own outputs. - **As a reward signal for RL without adversarial hardening.** The master-key result is exactly what an optimiser will find, and it will find it faster than you will. - **High-stakes criteria where your own annotators show low agreement.** A judge cannot be more reliable than the guideline it is imitating. - **Chasing small effect sizes.** A judge whose verdict flips when you swap the answer order cannot resolve a 2% regression. The position bias alone is larger than the effect you're measuring. That last one took me longest to accept. I spent weeks reporting on differences my instrument could not resolve, which is worse than reporting nothing. --- ### The shift I adopted a judge on the strength of a percentage and spent six weeks treating its output as a measurement. The percentage was real. What I inferred from it was not, and the gap between those two things had a published number on it the whole time. - **Correct for chance, then report per criterion.** Raw agreement is the number that flatters you; the deflation gap runs 33 to 41 points. - **Judge both orders, with a reference, on binary criteria.** Three cheap changes, and the reference-guided one alone moved a published failure rate from 70% to 15%. - **Attach a canary set and check it first.** A judge that rewards a lone colon invalidates everything else you measured that day. Take 100 items you already have human labels for, run your judge over each one twice with the answer order swapped, and compute both raw agreement and Cohen's kappa per criterion. > The distance between those two columns is the number nobody puts in a blog post, and it's the one that tells you whether your evaluation has been measuring anything at all. --- Source: https://himanshuat.com/blogs/llm-as-judge-what-the-numbers-actually-say ======================================================================== --- title: "Loop engineering" description: "Agents rarely fail because the model can't do the task. They fail because the same task succeeds on Monday and fails on Thursday while somebody is watching, and nothing in the loop notices. Fixing that meant engineering the loop instead of the prompt." date: "June 02, 2026" url: "https://himanshuat.com/blogs/loop-engineering" --- # Loop engineering The agent worked when I demoed it. A week later it failed in front of people, on a task it had already completed correctly twice. Nobody grades you on your average when they are watching the run happen. For two months after that I fixed the wrong layer. I rewrote the system prompt every time the agent did something stupid, and every rewrite fixed the failing case while quietly breaking a different one. > What I was short on was never prompt quality or model capability. It was consistency, and consistency is a property of the loop. ### Chapter 0 — What a loop actually is Anthropic describes the agent loop as four stages: **gather context, take action, verify work, repeat.** That's the atom. One agent improving one thing on repeat. The stage most homegrown loops skip entirely is verification, and it's the one that converts a stochastic process into something that trends toward correct. Skip it and you don't have a loop. You have a straight line that happens to run more than once. --- ### Step 1 — Measure `pass^k`, not `pass@1` The number that finally reframed this for me comes from tau-bench, the tool-agent-user benchmark Sierra published in June 2024. Alongside the usual `pass@1` they report `pass^k`: the fraction of tasks where **all k** independent attempts succeed. Not "at least one". All of them. State-of-the-art function-calling agents of the gpt-4o class solve under 50% of tasks on a single attempt. Run the same task eight times and require every run to succeed, and `pass^8` in the retail domain drops under 25%. Same model, same task, same prompt. The only variable is that you ran it more than once. That gap is the entire argument for treating the loop as the unit of engineering. A demo that works once tells you almost nothing about whether the system works. If you have ever committed to a date on the strength of a good screen recording, that number should sting a little. I had done exactly that, and the failing week was the bill arriving. > Prompt iteration optimises the mean. What kills you in production is the variance, and variance is a property of the loop, not of the string you pass in. **The tell:** if you have never run one task eight times in a row, you do not know your pass rate. You know your best case. --- ### Step 2 — Build the verification ladder, cheapest rung first Anthropic names three mechanisms for verification. Rules-based feedback, which is linting, type checking, and test runs. Visual feedback, meaning screenshots the model can look at. And LLM-as-judge for the fuzzy rules that no linter expresses. Their own caveat on the third is worth quoting because most people skip it: it carries latency tradeoffs and "isn't considered very robust". So the ladder has an order. **Deterministic checks first, judge last, and only for things a deterministic check genuinely cannot express.** There's a nice second-order point in their guidance about rules-based feedback: TypeScript with linting gives an agent more feedback layers than raw JavaScript does. The language choice is a loop design decision. More places for the environment to say "no, that's wrong" means more chances for the agent to correct itself without a human in the path. ```python # loop.py from dataclasses import dataclass, field from typing import Callable, Literal Verdict = Literal["pass", "fail", "unknown"] @dataclass class Check: """One rung on the verification ladder. `cost` is a rough ordering hint, not a measurement. Cheap deterministic checks run first so the expensive fuzzy ones only see work that already survived the mechanical ones. """ name: str cost: int run: Callable[[str], Verdict] @dataclass class Loop: """Gather context, take action, verify, repeat.""" checks: list[Check] = field(default_factory=list) def verify(self, artifact: str) -> tuple[Verdict, list[str]]: """Run checks cheapest-first and stop at the first hard failure. Returns the verdict plus every check name that ran, so the failure that stopped the ladder is recoverable from the trace. """ ran: list[str] = [] for check in sorted(self.checks, key=lambda c: c.cost): ran.append(check.name) verdict = check.run(artifact) if verdict == "fail": return "fail", ran return "pass", ran def step(self, context: str) -> tuple[str, Verdict]: """One full turn of the loop.""" artifact = self.act(context) verdict, _ = self.verify(artifact) return artifact, verdict ``` The ordering matters for cost as much as for quality. A judge call on an artifact that would have failed `tsc` is money you set on fire. **The tell:** if your most expensive check runs first, you are paying to grade work that a free check would have rejected. --- ### Step 3 — Treat context as a budget, not a window Anthropic frames context engineering as "the set of strategies for curating and maintaining the optimal set of tokens during LLM inference". The operative idea is the attention budget: it's finite, it has diminishing marginal returns, and the goal is "the smallest possible set of high-signal tokens". A million-token window is not a million tokens of usable attention. Chroma's context rot report from July 2025 is the measurement behind that. They tested 18 models, including Claude Opus 4 and Sonnet 4, o3, GPT-4.1, Gemini 2.5 Pro, and Qwen3, and found performance degrades consistently as input length grows **even on trivial tasks**. Not hard reasoning tasks. Trivial ones. When the semantic similarity between the needle and the question is low, degradation is faster. A single distractor in the haystack hurts, and four distractors compound the damage. The result I keep coming back to is the strangest one. Models did better on shuffled haystacks than on logically structured ones, and this held across all 18 models. Coherent surrounding text hurts retrieval compared with incoherent surrounding text. *I don't have a clean mechanistic story for why, and I'm suspicious of anyone who offers one confidently.* But it does undercut the intuition that neatly organised context is automatically better context. The cleanest argument for compaction I've seen is from LongMemEval, where a focused prompt of roughly 300 tokens was compared against a full prompt of roughly 113K tokens, with a large gap in favour of the focused one. > Three hundred tokens against a hundred and thirteen thousand. That is the whole case for curation in one comparison. Compaction in practice means summarising the conversation as it approaches the window limit and reinitialising with the summary. What Claude Code preserves through that is instructive: architectural decisions, unresolved bugs, and implementation details, while discarding redundant tool outputs. The reinitialised context is the summary plus the five most recently accessed files. ```python # compaction.py from dataclasses import dataclass KEEP_RECENT_FILES = 5 @dataclass class Transcript: """A conversation plus the files it touched, in access order.""" messages: list[dict] files_touched: list[str] def compact(transcript: Transcript, summarise) -> Transcript: """Summarise a near-full transcript and reinitialise from the summary. Recall first: the summariser is instructed to over-include. Precision comes later, by iterating on the prompt once you can see what a dropped detail actually costs downstream. """ summary = summarise( transcript.messages, keep=["architectural decisions", "unresolved bugs", "implementation details"], drop=["redundant tool outputs"], ) recent = transcript.files_touched[-KEEP_RECENT_FILES:] return Transcript( messages=[{"role": "user", "content": summary}], files_touched=recent, ) def clear_tool_results(messages: list[dict], keep_last: int) -> list[dict]: """The lightest-touch alternative to full compaction. Blanking stale tool results reclaims most of the space with none of the summarisation risk, because nothing gets paraphrased. """ seen = 0 out = [] for msg in reversed(messages): if msg.get("role") == "tool": seen += 1 if seen > keep_last: msg = {**msg, "content": "[cleared]"} out.append(msg) return list(reversed(out)) ``` The tuning rule Anthropic gives for compaction prompts is the one I'd have gotten backwards on my own: **maximise recall first, then iterate toward precision.** Dropping something important is much more expensive than carrying something redundant, so start greedy and trim later. Before you reach for compaction at all, tool result clearing is the lightest-touch version of the same idea, since it reclaims space without paraphrasing anything. Sub-agents are the other lever. A sub-agent can burn tens of thousands of tokens exploring and return a distilled summary of often 1,000 to 2,000 tokens. The search happened, the parent just never has to carry the transcript of it. **The tell:** if your agent's context grows monotonically across a run, you have a window, not a budget. --- ### Step 4 — Cap the run with mechanisms that already ship Every agent framework ships a turn cap, and every one of them is a blunt instrument that you still need. The OpenAI Agents SDK has `max_turns`, which raises `MaxTurnsExceeded`. LangGraph has `recursion_limit`, defaulting to 1000 super-steps, which raises `GraphRecursionError`, plus the managed `RemainingSteps` channel that lets a node see its remaining budget and wrap up deliberately rather than dying at the wall. Guardrail tripwires act as out-of-band terminators, killing a run on a condition that has nothing to do with turn count. There's an implicit budget most people never think about: **tool output size.** Claude Code truncates tool responses at 25,000 tokens by default, with pagination and filtering as the recommended pattern rather than dumping everything and hoping. If your tool can return a hundred thousand tokens of log, the truncation boundary is now part of your loop's semantics whether you designed it or not. Verbosity is a knob too. Anthropic's Slack example shows a `ResponseFormat` enum where the concise setting produced 72 tokens against the detailed setting's 206, roughly a third. Exposing that as a parameter the agent sets per call, rather than a global default, is close to free. ```python # budget.py from dataclasses import dataclass from enum import Enum TOOL_OUTPUT_LIMIT = 25_000 class ResponseFormat(str, Enum): """Verbosity as an explicit, per-call parameter.""" CONCISE = "concise" DETAILED = "detailed" @dataclass class Budget: """Turn and token caps, tracked together. Turns alone are a bad proxy: one turn that reads four files costs far more than one turn that writes a line of code. """ max_turns: int max_tokens: int turns_used: int = 0 tokens_used: int = 0 @property def remaining_turns(self) -> int: return max(0, self.max_turns - self.turns_used) def spend(self, tokens: int) -> None: """Record one turn. Raises when either cap is exhausted.""" self.turns_used += 1 self.tokens_used += tokens if self.turns_used >= self.max_turns: raise BudgetExhausted(f"turn cap {self.max_turns} reached") if self.tokens_used >= self.max_tokens: raise BudgetExhausted(f"token cap {self.max_tokens} reached") def wrap_up_signal(self) -> str | None: """Tell the model to land the plane while it still has room to.""" if self.remaining_turns <= 2: return "Two turns left. Summarise state and stop." return None class BudgetExhausted(RuntimeError): """Raised when a run hits a cap. Never caught silently.""" ``` **The tell:** if you have ever raised `max_turns` because a run died at the cap, check whether the extra turns went anywhere. Mine never did. --- ### Step 5 — Anchor the loop to something outside the model For work that spans multiple context windows, the scaffolding around the model is called a harness, and the important finding is that **compaction alone is not sufficient.** Summarising well still leaves you with an agent that has no durable record of what it committed to. Two failure modes show up by name: - **Over-ambition.** Agents try to one-shot the whole app, exhaust context mid-implementation, and leave behind half-finished features that nothing documents. - **Premature completion.** A later instance, reading a summary rather than the work, declares the job done. Both got fixed by the same thing, and it's mundane. A structured external checklist, held as a JSON feature list with more than 200 structured end-to-end test cases, each marked `passes: false` until proven otherwise, plus an explicit instruction that removing or editing tests is unacceptable. That last clause is not paranoia. An agent optimising for a green checklist will absolutely delete the failing test if you don't forbid it. ```python # harness.py import json from pathlib import Path CHECKLIST = Path("features.json") def load_checklist() -> list[dict]: """Read the durable feature list that survives every context window.""" return json.loads(CHECKLIST.read_text()) def next_unfinished(checklist: list[dict]) -> dict | None: """Pick the next unproven feature, in declaration order. Declaration order, not model preference. Letting the agent choose its own next task is how you get four half-built features instead of one finished one. """ for feature in checklist: if not feature["passes"]: return feature return None def mark_pass(checklist: list[dict], feature_id: str, evidence: str) -> list[dict]: """Flip a feature to passing, but only with attached evidence. Evidence means a browser transcript, not a claim. Code inspection is explicitly not acceptable proof that a feature works. """ if not evidence.strip(): raise ValueError(f"{feature_id}: cannot pass without evidence") for feature in checklist: if feature["id"] == feature_id: feature["passes"] = True feature["evidence"] = evidence CHECKLIST.write_text(json.dumps(checklist, indent=2)) return checklist ``` The verification detail from that same work is the one I'd tattoo somewhere. Claude initially marked features complete on the basis of code inspection alone. It read the code, the code looked right, the box got ticked. The fix was an explicit instruction to drive a browser and test the feature the way a human user would. > "I read the implementation and it looks correct" is not evidence, from an agent or from anyone else. **The tell:** if the only record of what the agent committed to lives in the agent's context, it will forget it and then tell you it's done. --- ### What breaks it **The loop grades its own homework.** Everything in Step 2 exists because a model checking its own output goes easy on itself. The ladder works only if the bottom rungs are things the model cannot argue with — a test that ran, a type that checked. A judge reading the same context that produced the artifact is not verification. **The metric drifts and the loop can't see it.** A loop only sees its own metric. It cannot ask whether the target is right. A support bot tuned on ticket resolution rate learns to close tickets fast rather than solve them, and the number climbs for months while satisfaction drops. That's Goodhart's law with a scheduler attached. **The no-progress detector, which you should treat as an opinion.** Here's where I have to be honest about the state of the evidence. I needed a no-progress detector, because a run that spins burns real money and produces nothing anyone can use, and I needed it that week. I built one. It hashes the observable state after each turn, tracks repeated tool-call n-grams, and trips when the last several turns are structurally identical to earlier ones. It works on the traces I had, and I shipped it because the alternative was letting runs spin. *It is also entirely my own engineering judgment.* I could not find a primary source publishing a specific, validated no-progress or oscillation detection algorithm — not state-hash repetition, not action edit distance, not n-gram repetition of tool calls. So I'm not going to dress mine up as established practice or put a number on how well it works. What is documented, and what I'd build on before building a detector, are turn caps, `RemainingSteps`, and an external checklist that makes progress observable from outside the model's own account of itself. A detector guesses at progress. A checklist measures it. If you want a map of how these things fail rather than a list of anecdotes, MAST is the one to read. It's a taxonomy from Berkeley of 14 failure modes across 3 categories — system design, inter-agent misalignment, and task verification. It was built from more than 1,600 annotated traces across 7 frameworks, and the human annotators reached a Cohen's kappa of 0.88, which is a level of agreement that makes the categories worth taking seriously rather than treating as one team's vocabulary. --- ### Errors are the teaching signal, so write them like documentation An agent reads your error messages more carefully than your users ever will, and unlike your users it cannot go ask someone. Anthropic's guidance is that error messages should communicate specific, actionable improvements and include correctly-formatted examples. A tool that returns `ValidationError: invalid input` teaches nothing. A tool that returns the failing field, the expected shape, and a working example changes the next attempt. Two smaller details from the same guidance changed my tool design: - **Semantic identifiers beat UUIDs**, and the claim is specific: they significantly improve Claude's precision in retrieval tasks by reducing hallucinations. A model can pattern-match `invoice_2026_04_acme` in a way it simply cannot do with a hex blob. - **Tool namespacing has non-trivial effects** depending on whether you prefix or suffix, which sounds like superstition until you watch a model consistently reach for the wrong one of two similarly named tools. --- ### When loop engineering is the wrong choice Most tasks don't need any of this. Reaching for the machinery when you don't need it burns money and adds ways to fail. - **The task runs once and a human reads the output.** A one-off script, a migration you'll watch. The verification ladder is overhead when you *are* the verification ladder. - **You don't know what correct looks like yet.** You cannot build a check for a target you haven't defined. Explore first with a single steerable agent, then engineer the loop once the success condition is stable. - **The failure mode is capability, not consistency.** If the agent fails all eight runs rather than three of eight, the loop isn't your problem. No amount of verification turns a task the model can't do into one it can. - **There's no cheap deterministic check available.** If every rung of your ladder is an LLM judge, you have built an expensive committee, not a loop. Find a real signal first or accept that this stays human-verified. **The tell is Step 1.** Run the task eight times. If it succeeds eight times, you don't have a consistency problem. If it succeeds zero times, you don't have a loop problem. Everything in between is what this post is for. --- ### The shift I spent two months tuning prompts and about two weeks building loop machinery, and the two weeks produced almost all of the reliability. That ratio was backwards, and I'd have caught it before I made any promises if I'd measured `pass^k` instead of running the happy path once the night before a demo and feeling good about it. The other thing I got wrong: I raised `max_turns` whenever a run died at the cap. Every single time it felt like the right move, and every single time the extra turns went into the same spin. More turns is not better. Without a progress signal, raising the cap converts a failure into an expensive failure, and you pay for the privilege of watching it fail later. - **Measure the variance, not the mean.** `pass^k`, not `pass@1`. - **Put the cheap deterministic check at the bottom of the ladder.** The judge is the last rung, not the first. - **Anchor progress outside the model.** A checklist it isn't allowed to edit beats a detector that guesses. The concrete next step, if you take one thing from this: pick your three most important tasks, run each one eight times against your current agent, and count how many succeed all eight times. > That number is your real pass rate. Everything in this post is a response to what that number looks like the first time you compute it. --- Source: https://himanshuat.com/blogs/loop-engineering ======================================================================== --- title: "LoRA finetuning: the hyperparameters settled in a week, the data never did" description: "An early draping adapter kept painting the same flat plane across the back of the shoulder, on prompts with nothing in common. That is memorisation, not physics, and it traced back to one shoot. The config for these LoRAs settled after a handful of runs and never moved again. Everything that improved after that came out of the dataset." date: "April 08, 2025" url: "https://himanshuat.com/blogs/lora-finetuning-what-moved-the-needle" --- # LoRA finetuning: the hyperparameters settled in a week, the data never did An early Draping expert kept producing the same flat plane across the back of the shoulder. Different prompts, nothing in common between them, same plane every time. That's the signature of memorisation, not of a learned prior. It traced back to a cluster of images from one shoot where the fabric had been pulled tight and pinned out of frame. I went into LoRA training expecting a long tuning phase. Rank sweeps, learning rate schedules, the whole ritual. What actually happened is that the config settled after a handful of runs and then sat unchanged for months while I rebuilt the dataset more than once. > Rank and alpha have a real effect and it's a small, bounded one. The dataset has an unbounded one. ### Chapter 0 — What alpha over rank actually controls A LoRA doesn't rewrite the base weights. It learns a low-rank product and adds it, scaled: `ΔW = (α / r) · B A` Rank is capacity. Alpha over rank is how loudly that capacity speaks. Two knobs, and they do different jobs. Both experts here use the same settings. Expert A is Draping Physics, trained on sarees, hanfus, kimonos and complex haute couture. Expert B is Occlusion and Depth, trained on hands on hips, arms crossed, and jewelry sitting over clothing. Same rank, same alpha, same schedule, different data. That symmetry is deliberate. If two adapters are going to be merged into one backbone later, you want their update magnitudes comparable rather than accidentally scaled apart by their training configs. --- ### Step 1 — Set the scale below one, then stop touching it **Rank 32, alpha 16.** The ratio matters more than either number alone. At alpha 16 over rank 32 the scale is 0.5. | alpha | scale | what it did | |---|---|---| | 16 | 0.5 | shipped. drape present, base-model look intact | | 32 | 1.0 | drape more emphatic, images more obviously overfit | | 64 | 2.0 | common convention. I never ran it | Alpha equal to rank and alpha at twice the rank are both common conventions. Running at 0.5 damps every update the adapter makes to the base weights. The practical effect is roughly the same as halving the learning rate for the adapter's contribution while leaving the optimizer's own step size alone. That's the behaviour I wanted. These adapters exist to add a physics prior to a model that's already very good at images. If the adapter shouts, you get the thing everyone gets on their first LoRA: the output stops looking like the base model and starts looking like the training set, including its backgrounds, its colour grading, and its favourite camera angle. Rank 32 gives enough capacity to represent something as structured as pleat geometry. The low alpha keeps that capacity from bleeding into everything else. **Learning rate 1e-4 with cosine annealing.** Standard, and it stayed standard. Cosine matters more than the peak value here. The last stretch of training on a small, visually consistent dataset is where an adapter picks up the dataset's incidental style, and annealing down keeps those late steps small. **The tell:** if raising alpha makes the drape more emphatic and the images more obviously overfit at the same time, you have found the ratio's real job. It is not a quality knob. --- ### Step 2 — Freeze the text encoder and stay on the attention projections **UNet only, attention projections only.** The frozen text encoder is a decision about what kind of adapter this is. Training the text encoder too lets the model redefine words. That's right when you're teaching a new token: make this made-up word mean this person, or this illustration style. Draping physics isn't a word-meaning problem. The base model knows what a saree is. What it doesn't know is what the fabric does under gravity, and that lives in the denoiser's spatial attention rather than in a text embedding. Leaving the text encoder trainable also hands the optimizer a cheaper route to the same loss. It can rebind the caption tokens to the training set's overall look, reconstruct better, and learn nothing about folds. Freezing it removes that option. The same reasoning keeps the adapter on the attention projections rather than every linear layer in the network. Query, key, value and output projections are where spatial relationships between regions get decided, and a fold is a spatial relationship. **The tell:** ask whether the base model already knows the noun. If it does, the missing knowledge is spatial, and the text encoder is not where it lives. --- ### Step 3 — Size the batch around the buckets, not the VRAM **Batch size 4 with gradient accumulation.** There's room on a 48GB L40S to raise the real batch instead, and I deliberately didn't, because of bucketing. Images of different shapes can't be stacked into one tensor, so a physical batch is drawn from inside a single aspect bucket. In this dataset a bucket is effectively a shot type: tall full-length shots in one, wide flat-lays in another. A large physical batch means a long run of gradient signal from one shot type. Accumulating small batches lets the effective batch span buckets, so a step averages across shapes rather than specialising to whichever bucket it landed in. It also keeps the number of optimizer steps per epoch high enough that a cosine schedule has something to anneal over, which a few thousand images at a large true batch would not. **AdamW8bit.** Eight-bit optimizer states. The memory it frees goes into resolution rather than batch size, which for garment work is the better place to spend it. Fabric detail is the entire point, and a pleat rendered at low resolution in training is a pleat the adapter never learns. That's the whole config: ```toml # configs/drape_lora.toml # Expert A: Draping Physics. Kohya-ss network trainer. # Expert B (occlusion) uses this file with the dataset paths swapped. [model] pretrained_model_name_or_path = "/models/flux/base" vae = "/models/flux/vae" mixed_precision = "bf16" save_precision = "fp16" [network] network_module = "networks.lora_flux" network_dim = 32 # rank network_alpha = 16 # scale = alpha / dim = 0.5 network_train_unet_only = true # text encoder stays frozen [optimizer] optimizer_type = "AdamW8bit" learning_rate = 1e-4 lr_scheduler = "cosine" lr_warmup_steps = 100 max_grad_norm = 1.0 [training] train_batch_size = 4 gradient_accumulation_steps = 4 gradient_checkpointing = true cache_latents = true cache_latents_to_disk = true seed = 42 max_train_epochs = 12 save_every_n_epochs = 1 # keep every epoch; the best one is chosen by eye [dataset] resolution = "1024,1024" enable_bucket = true # sarees and full-body shots are not square bucket_reso_steps = 64 min_bucket_reso = 768 max_bucket_reso = 1536 caption_extension = ".txt" shuffle_caption = false # caption order is load-bearing, see below keep_tokens = 1 # the trigger token stays in position one [[dataset.subsets]] image_dir = "/data/drape/saree" num_repeats = 1 [[dataset.subsets]] image_dir = "/data/drape/hanfu_kimono" num_repeats = 1 [[dataset.subsets]] image_dir = "/data/drape/couture" num_repeats = 1 [logging] logging_dir = "/logs/drape" log_with = "tensorboard" ``` Two lines in there are doing more work than they look like they are. `shuffle_caption = false` with `keep_tokens = 1` pins the trigger token to the front of every caption, which makes the adapter's activation predictable at inference instead of something that partially fires whenever a related word appears. And bucketing is not optional for this domain. A full-length saree shot is tall, a flat-lay couture reference is wide, and if you square-crop everything to fit a fixed resolution you cut off precisely the pallu drape you're trying to teach. `save_every_n_epochs = 1` is there because no validation loss on this task correlates with "the pleats look right". Epoch selection stays a human judgement, so every epoch has to survive to be judged. **The tell:** if raising the physical batch would leave the cosine schedule too few optimizer steps to anneal over, the VRAM headroom is not for the batch. Spend it on resolution. --- ### The invocation One node, one config file, and a guard before anything expensive starts. ```bash # scripts/train_drape_lora.sh # Single L40S node (48GB). Run from the kohya-ss/sd-scripts checkout. set -euo pipefail RUN_NAME="drape_r32a16_$(date +%Y%m%d_%H%M)" DATA_ROOT="/data/drape" OUT_DIR="/artifacts/lora/${RUN_NAME}" mkdir -p "${OUT_DIR}" # Fail here rather than deep into a run: every image needs a caption file. missing=$(find "${DATA_ROOT}" -name '*.jpg' | while read -r img; do [ -f "${img%.jpg}.txt" ] || echo "${img}" done | wc -l) if [ "${missing}" -gt 0 ]; then echo "aborting: ${missing} images have no caption file" >&2 exit 1 fi accelerate launch \ --num_cpu_threads_per_process 8 \ --mixed_precision bf16 \ flux_train_network.py \ --config_file configs/drape_lora.toml \ --output_dir "${OUT_DIR}" \ --output_name "${RUN_NAME}" \ --save_state # Every epoch is kept. Selection happens by generating the fixed evaluation # prompt set against each checkpoint and looking at the results. ls -1 "${OUT_DIR}"/*.safetensors echo "checkpoints written to ${OUT_DIR}" ``` The caption check at the top exists because I lost a run to it. I'm not going to quote wall-clock times, cost, or a loss curve here. I don't have numbers from those runs that I'd stand behind publishing, and a training curve on this task wouldn't tell you much anyway, since the loss goes down smoothly through epochs that are visibly getting worse. --- ### Step 4 — Filter mechanically first, then go hunting for what passed Once the config was fixed, every real improvement came from the dataset. Filtering came before all of it. Public images go through a mechanical pass first: resolution, sharpness, aspect ratio. This throws away a lot and it should. An image that's soft at the garment boundary teaches the adapter that garment boundaries are soft, and there's no later stage that undoes that. ```python # data/filter_corpus.py """First-pass technical filter over a candidate image corpus. Nothing here is about whether an image is *good*. It's about whether the image is technically usable as a training example for fabric structure. """ from dataclasses import dataclass from pathlib import Path import cv2 import numpy as np @dataclass(frozen=True) class FilterConfig: min_short_side: int = 1024 # below this, fabric texture is already gone min_sharpness: float = 120.0 # variance of Laplacian min_aspect: float = 0.5 # portrait limit max_aspect: float = 1.6 # wide-crop limit max_clipped_fraction: float = 0.04 # blown highlights eat white fabric def sharpness(gray: np.ndarray) -> float: """Variance of the Laplacian. Crude, fast, and good enough at scale.""" return float(cv2.Laplacian(gray, cv2.CV_64F).var()) def clipped_fraction(gray: np.ndarray) -> float: """Fraction of pixels at or near the top of the range.""" return float((gray >= 250).mean()) def inspect(path: Path, cfg: FilterConfig): """Return (accepted, reason). Reason is None when accepted.""" image = cv2.imread(str(path)) if image is None: return False, "unreadable" h, w = image.shape[:2] gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) aspect = w / h if min(h, w) < cfg.min_short_side: return False, f"resolution {w}x{h}" if not cfg.min_aspect <= aspect <= cfg.max_aspect: return False, f"aspect {aspect:.2f}" if sharpness(gray) < cfg.min_sharpness: return False, "soft" if clipped_fraction(gray) > cfg.max_clipped_fraction: return False, "clipped highlights" return True, None def run(root: Path, cfg: FilterConfig = FilterConfig()): kept, rejected = [], {} for path in sorted(root.rglob("*.jpg")): ok, reason = inspect(path, cfg) if ok: kept.append(path) else: rejected[reason] = rejected.get(reason, 0) + 1 print(f"kept {len(kept)} of {len(kept) + sum(rejected.values())}") for reason, count in sorted(rejected.items(), key=lambda kv: -kv[1]): print(f" rejected {count:>6} {reason}") return kept ``` What that filter can't catch is the sample type that hurt most: catalogue shots where the drape is fake. Studio garments get clipped, pinned, or taped at the back so the front hangs the way the stylist wants. Those images are sharp, well lit, high resolution, and they pass every check in that script, while being photographs of fabric under a tension gravity never applied. It's the exact opposite of what a draping adapter should learn. The flat shoulder plane is how I found it, by accident. The diagnostic that came out of it has stayed useful, and it forks on one question: ```mermaid flowchart TD A[repeated artefact] --> B{style or geometry?} B -->|style| C[alpha too high] B -->|geometry| D[find the cluster] C --> E[retrain] D --> F[rebuild dataset] F --> E ``` **The tell:** repeated *style* across unrelated prompts is an alpha problem. Repeated *geometry* is almost always a specific pile of images you need to go find. --- ### Step 5 — Caption what varies, and vary the relationship, not the pose Captions have to describe what varies and stay silent about what doesn't. If every saree image says "saree" and every one of them also happens to be shot against a grey backdrop, and the caption never mentions the backdrop, the adapter learns that grey backdrop is part of what "saree" means. Writing captions that name the incidental properties is how you tell the model those properties are separable. This was more work than all the hyperparameter tuning combined and it produced more improvement. The same principle governs what you photograph. The Occlusion and Depth expert is trained on hand-garment interaction, and the first version of that dataset had a lot of near-duplicates: the same hand-on-hip pose from slightly different angles. That teaches a pose, not a relationship. What the adapter needs is the same *relationship*, a hand in front of fabric, across genuinely different hands, garments, and fabric behaviours behind the hand. **The tell:** read your captions and ask what they never mention. Every property shared by every image and named by no caption is now part of what your trigger token means. --- ### Step 6 — Pick the epoch on a frozen grid, at partial strength, outside the domain There was no automated metric for drape quality at the time, and "we looked at the outputs" is only half the story of what replaced it. Every saved epoch gets generated against the same frozen prompt set with the same fixed seeds. Frozen and fixed are the load-bearing words: change the prompts and you're comparing two things at once, change the seeds and you're mostly looking at sampling noise. The result is a grid, prompts down one axis and epochs across the other. The eye reads that grid well even though it scores any single image badly. What you're watching for is two failures moving in opposite directions. | symptom | where it shows | reading | |---|---|---| | mushy pleats, pallu won't commit to a shoulder | on the garment | undercooked | | backgrounds converging, colour grading drifting, faces less varied | nowhere near the garment | overcooked | You want the last epoch where drape improved and the second kind of drift hadn't started. Two other checks mattered more than they sound. Look at the adapter at reduced strength as well as at full weight, because it will be merged at a coefficient below one alongside the occlusion expert, and an adapter whose drape only appears at full strength is useless to that merge. I threw one away that looked excellent in isolation for exactly this reason. Then generate things well outside the training domain. A plain t-shirt, a jacket, anything the Draping expert has no business touching. Catastrophic forgetting announces itself outside the domain first, and a checkpoint that renders a beautiful saree and a broken t-shirt has learned the dataset rather than the physics. **The tell:** if the epoch you picked is the final one you trained, check again. The one you want is usually earlier. --- ### Four things to build before the next run - **The evaluation prompt set, before the first training run.** Epoch selection is subjective, and a frozen prompt set at least makes it consistently subjective. We ended up with one, and having it from the start would have saved the early runs that I now can't compare against anything. - **Dataset versioning as strict as code versioning.** Half of my confusion about why one adapter behaved differently from another came down to not knowing exactly which images went into it. - **An explicit test that separates "learned the concept" from "learned the dataset."** That distinction is the actual outcome of the alpha-to-rank choice, and it deserved a deliberate test rather than my eye on a grid of samples. - **The alpha ablation.** Hold the dataset fixed, train at 0.5, 1.0, and 2.0 scale, and evaluate for style contamination rather than for drape quality. I picked 0.5 by reasoning and confirmed it by looking, which isn't the same as measuring it. --- ### What breaks it **An image with no caption file.** Kohya will happily train on an empty caption, and the result is an adapter that has learned to associate the trigger with nothing in particular. Nothing crashes. The run completes, the checkpoints look normal, and the adapter is quietly worse. That's why the guard sits at the top of the launch script rather than being a thing I remember to check. **The loss curve.** It goes down smoothly through epochs that are visibly getting worse, which means there is nothing here to early-stop on. No validation loss on this task correlates with "the pleats look right", so a metric that keeps improving is not evidence and shouldn't be read as any. **Tuning instead of rebuilding.** The cheapest debugging tool was accepting that a dataset was wrong and throwing the run away, rather than trying to rescue it with a different learning rate. I resisted this for longer than I should have because retraining feels expensive and tuning feels productive. Both feelings were wrong in the same direction. --- ### When this setup is the wrong choice The config above is shaped for one job: adding a physical prior to a base model that already renders the subject well. Point it at a different job and most of these decisions invert. - **You're teaching a name, not a behaviour.** Make this made-up word mean this person, this illustration style. That is a word-meaning problem, and freezing the text encoder removes the exact capacity you need. - **The base model can't render the subject at all.** Everything here assumes the model already knows what a saree is and only needs to learn what the fabric does. Nothing in these runs is evidence that a damped rank-32 adapter can teach a subject from nothing. - **Nobody can tell good from bad by eye, and you have no metric.** Epoch selection here is a human looking at a grid. If that judgement isn't available and no automated metric exists either, you have a training procedure with no selection procedure attached, which is not a procedure. - **The dataset is fixed and can't be rebuilt.** The config settled in a week and the data never did. If you can't re-filter, re-caption, or re-shoot, the part of this that produced the improvement isn't available to you, and no learning rate substitutes for it. --- ### The shift I expected the interesting surface to be the optimizer settings. It wasn't. Rank and alpha have a real effect and it's a small, bounded one, and I got to the end of it in a handful of runs. Everything after that was a data question wearing a training question's clothes. The flat shoulder plane looked like an overfitting problem and was a folder of pinned fabric. The adapter that only worked at full strength looked like a good checkpoint and was an unusable one. - **Freeze the config early and treat every later question as a data question.** The ratio is a design decision, not a dial you keep turning. - **Filter for what a photograph can't fake, not just for what it renders sharply.** Sharp, well-lit, high-resolution images of fabric under artificial tension pass every mechanical check you can write. - **Judge the checkpoint at partial strength and outside its domain.** Full-strength, in-domain samples are the ones that flatter it. The concrete next step, if you take one thing from this: before your next run, write down the evaluation prompt set and the seeds, and don't change either of them again. > An adapter you selected by eye on a frozen grid is subjective. An adapter you selected by eye on a grid you kept editing isn't even that. --- Source: https://himanshuat.com/blogs/lora-finetuning-what-moved-the-needle ======================================================================== --- title: "Setting up Linters, Syntax Highlighter, NerdTree, and More" description: "A guide on configuring Neovim as a full-blown IDE." date: "September 12, 2023" url: "https://himanshuat.com/blogs/my-nvim-setup" --- # Setting up Linters, Syntax Highlighter, NerdTree, and More Before we begin, a quick word on why Vim is still worth learning in 2022, and a bit about how I got into it. After that, a practical guide to configuring your Neovim into a full-blown IDE. ## Why do People use Vim in 2022? There is a compelling reason why Vim has remained relevant to most developers for nearly half a century. Among developers, the functionality that Vim offers to edit multiple files with a few keystrokes is simply unmatched. Being able to progress on your deliverables and explore candidate solutions at the speed of thought is crucial among developers working in mid to large codebases. That is the reason I got into Vim. The learning curve is quite steep. It took me a week to get comfortable with the keystrokes, but it pays off quite well. (I found this cheat sheet helpful: [Vim Cheat Sheet](https://vim.rtorr.com/).) Watching more experienced developers like ThePrimeagen, George Hotz, and Jason Turner move through multiple files so fluidly, reimplementing their design and testing their code, made me want to incorporate Vim into my personal workflow. It allowed me to focus more on the code instead of spending too much time on popups in a GUI editor. Although I don't write on Vim (and Neovim) every time, text editors and IDEs offer Vim keybindings, minimizing the lag you spend clicking on the GUI. The point is, no matter where you go in your dev journey, you should give Vim a try. ## Vim and Neovim Neovim is a result of the Vim community's effort to support better scripting through Lua, modern plugins, and integration with modern GUIs (Wikipedia Contributors, 2022). The Neovim project started in 2014 with its main design focus dedicated to the extensibility and maintainability of Vim. Most commands and Vim configurations work in Neovim by default as the Neovim project is a forked and more feature-rich version of Vim. The title of this section naturally lends itself to the topic of which is better, but I leave that to your judgment to decide which works better for you. Personally, I find Neovim more friendly as code completion and static analyzers are simpler to configure. ## Configuring Neovim All configurations for Neovim are stored in the `~/.config/nvim` directory. But unlike Vim which is available in most Linux distributions, you have to make sure that you have Neovim by running `nvim -v`. If it doesn't output the version, you have to install it using: ```bash sudo apt install neovim ``` (I'm using Ubuntu; your distro may have a different package manager to install Neovim.) Next, we have to install Git and Curl to set up Vim Plug, which allows you to install Vim plugins. But first, let's make sure that we are up to date. ```bash sudo apt update sudo apt install git curl ``` To install Vim Plug, run the shell command for Linux: ```bash sh -c 'curl -fLo "${XDG_DATA_HOME:-$HOME/.local/share}"/nvim/site/autoload/plug.vim --create-dirs \ https://raw.githubusercontent.com/junegunn/vim-plug/master/plug.vim' ``` Now, we navigate to `~/.config/nvim`. If this directory is not available, make one using: ```bash cd ~/ && mkdir -p .config/nvim ``` We put all our user-defined configurations in the `~/.config/nvim/init.vim` file. Don't panic if the blank screen opens. The first keys you should know to edit files in Vim are: - `i` for insert - `[esc] + q` for quit - `[esc] + w` for save - `[esc] + wq` for save and quit Here are some basic configurations: ```vim :set number " show line numbers :set relativenumber " show line number on the current line relative to other lines :set autoindent " sets newline to inherit the indentation of prev lines :set tabstop=4 " indents using 4 spaces :set shiftwidth=4 " sets 4 spaces indents when shifting :set smarttab " sets `tabstop` number of spaces when the tab is pressed :set softtabstop=4 " sets 4 spaces when tab or backspace is pressed :set mouse=a " enables the mouse for scrolling and resize ``` You may extend this initial configuration as you like. Refer to: [Top 50 Vim Configuration Options](https://www.shortcutfoo.com/blog/top-50-vim-configuration-options/). If you're satisfied, save and exit. ## Setting Up Plugins Your Vim plugins are declared between `call plug#begin()` and `call plug#end()`. Inside, we place a set of `Plug [plugin URL]` commands. And when we're satisfied, hit `[esc] + :PlugInstall`. To enable the tagbar for my configuration, use: ```bash sudo apt install exuberant-ctags ``` My personal Vim plugins are as follows: ```vim call plug#begin('~/.config/nvim/plugged') Plug 'http://github.com/tpope/vim-surround' " Surrounding ysw) Plug 'https://github.com/preservim/nerdtree' " NerdTree Plug 'https://github.com/tpope/vim-commentary' " For Commenting gcc & gc Plug 'https://github.com/vim-airline/vim-airline' " Status bar Plug 'https://github.com/ap/vim-css-color' " CSS Color Preview Plug 'https://github.com/rafi/awesome-vim-colorschemes' " Retro Scheme Plug 'https://github.com/ryanoasis/vim-devicons' " Developer Icons Plug 'https://github.com/tc50cal/vim-terminal' " Vim Terminal Plug 'https://github.com/preservim/tagbar' " Tagbar for code navigation Plug 'https://github.com/terryma/vim-multiple-cursors' " CTRL + N for multiple cursors call plug#end() nnoremap :NERDTreeFocus nnoremap :NERDTree nnoremap :NERDTreeToggle nmap :TagbarToggle let g:NERDTreeDirArrowExpandable="+" let g:NERDTreeDirArrowCollapsible="~" ``` NerdTree provides a graphical view of your file tree: ```jsx {`NerdTree ``` For the language server, code completion, and static analysis tools, we use `coc.vim` plugin. This requires NodeJS version 14+ and its package manager NPM. ```bash curl -sL https://deb.nodesource.com/setup_14.x | sudo -E bash - sudo apt-get install -y nodejs && sudo apt install npm ``` Now that we have the prerequisite for `coc.nvim`, open `init.vim` and add the following: ```vim call plug#begin() ... Plug 'https://github.com/neoclide/coc.nvim' " Auto Completion ... call plug#end() nnoremap :call CocActionAsync('jumpDefinition') ``` Head over to the `~/.config/nvim/plugged/coc.nvim` directory and build yarn with the following commands: ```bash sudo npm install -g yarn yarn install yarn build ``` To be able to configure the linters, we have to install different modules for different languages we're running on. For this example, we will set up Python3. (Make sure that Python3 and Pip3 are installed.) Install static analysis tool for Python (Jedi is one example but there are also others): ```bash pip3 install jedi ``` Inside our Python file, type: `[esc] + :CocInstall coc-python`. And there we have it! Thanks for getting this far and enjoy exploring tools in your terminal! ## References - Wikipedia contributors. (2022, August 26). Vim (text editor). In Wikipedia, The Free Encyclopedia. Retrieved 14:12, September 1, 2022, from [Wikipedia]() - Choi, J. (2022). Vim-Plug. [GitHub](https://github.com/junegunn/vim-plug). --- Source: https://himanshuat.com/blogs/my-nvim-setup ======================================================================== --- title: "One LoRA per failure mode" description: "Put a hand on a hip and our try-on pipeline painted the garment straight over the fingers. The obvious fix, one bigger dataset covering every hard case at once, quietly made the draping worse. Splitting the finetune into two adapters, one per named failure mode, is what actually worked." date: "May 20, 2025" url: "https://himanshuat.com/blogs/one-lora-per-failure-mode" --- # One LoRA per failure mode Put a person in the input image with one hand resting on their hip, run the pipeline, and the generated garment gets painted over the fingers. The hand is just gone, absorbed into the fabric. The bug had a name inside the team before it had a ticket. We called it the "Amputated Hand", and for a few weeks I tried to fix it the way everyone tries to fix it, by adding more of that kind of image to the finetune we already had. > "Hard cases" is not a category. It is two teachers writing opposite instructions into the same fixed adapter budget, and what you get back is the average. ### Chapter 0 — What one adapter can hold We started with one LoRA trained on "hard cases". That set was a pile of everything the base model got wrong: sarees, kimonos, hanfus, couture with unusual topology, plus the occlusion images that produced the amputated hands. One adapter, one training run, one number to tune. It's the cheapest thing to build and it's what I'd do again if I hadn't watched it fail. It failed because the two biggest problems in that pile want opposite behaviour from the model. Draping is a generation problem. A saree pallu has to fall over the shoulder with volume, the pleats have to gather at the waist and cast their own shadows, and the adapter's job is to push the model to *invent* structure inside the masked region that the base model has no prior for. The training signal rewards high frequency detail surviving the denoise. Occlusion is a suppression problem. When a hand sits on a hip, the correct behaviour inside part of the masked region is to generate nothing and leave the hand alone. The adapter's job is to teach the model that some pixels inside the mask are foreground and immutable. | | drape | occlusion | |---|---|---| | problem type | generation | suppression | | adapter's job | invent structure inside the mask | leave marked pixels alone | | training signal rewards | detail surviving the denoise | the mask being respected | At rank 32 the adapter has a fixed budget of directions it can write into the base weights. Asking it to simultaneously mean "add structure here" and "add nothing here" spends that budget on an average of the two. When I added occlusion images to the mixed set, the hands started surviving and the pleats went flat again. When I rebalanced toward drape, the hands came back off. I was tuning a dataset mixing ratio, and **every value of that ratio cost a full training run to evaluate.** There's a second problem with the single finetune that took longer to notice. When it regressed, I couldn't tell which images caused it. A mixed dataset gives you one loss curve and no attribution. --- ### Step 1 — Sort the rejects before you collect anything So I stopped collecting and started sorting. The honest starting state was a folder of a few hundred rejected PNGs with timestamps for filenames, and the reasons they'd been rejected living in review threads rather than anywhere structured. That's the condition most of these projects are in at the start. The tidy version of this story opens with a defect taxonomy nobody actually has. The first pass was a contact sheet and a person. Print the rejects in a grid, go through them, say out loud what's wrong with each one. That produced about a dozen phrases and the phrases repeated. "Hand gone." "Pleats flat." "Looks like a bathrobe." Nobody agreed on wording, which is why the classifier below matches substrings against a blob of text rather than reading a clean enum. The second pass was making the tagging cheap enough that it kept happening, so the reviewer's job became one line of free text with no dropdown and no schema. Then I ran this over the log. ```python # triage_failures.py """Bucket rejected try-on renders into named failure modes. Reviewers tag a rejected render with short free-text labels ("hand gone", "pleats flat", "looks like a bathrobe"). This maps those labels onto the failure modes we actually name, and prints the distribution, so that dataset collection targets a mode instead of a vague notion of "hard cases". """ import json from collections import Counter from pathlib import Path # Label fragments mapped to the failure mode they belong to. Deliberately not # exhaustive: anything unmatched lands in "unclassified" and gets read by hand, # which is how new modes get discovered. MODE_PATTERNS = { "drape": ["pleat", "fold", "flat texture", "pallu", "no volume", "plastered"], "occlusion": ["hand", "finger", "arm", "jewel", "necklace", "painted over"], "style_drift": ["wrong texture", "lost weave", "flat colour", "not the fabric"], "semantic": ["bathrobe", "wrong garment", "looks like a dress"], } def classify(tags: list[str]) -> set[str]: """Return every failure mode whose patterns appear in a render's tags. A render can belong to more than one mode. A saree rendered flat *and* over a hand is two separate defects and should be counted twice, otherwise the rarer mode gets hidden behind the common one. """ blob = " ".join(tags).lower() modes = {mode for mode, pats in MODE_PATTERNS.items() if any(p in blob for p in pats)} return modes or {"unclassified"} def summarise(reject_log: Path) -> Counter: counts = Counter() for line in reject_log.read_text().splitlines(): record = json.loads(line) for mode in classify(record["tags"]): counts[mode] += 1 return counts if __name__ == "__main__": counts = summarise(Path("data/rejects.jsonl")) total = sum(counts.values()) for mode, n in counts.most_common(): print(f"{mode:>14} {n:>5} {n / total:6.1%}") print(f"{'total defects':>14} {total:>5}") ``` The distribution isn't the interesting part. What mattered is that two buckets were big enough to justify their own training run and the other two weren't. That left drape and occlusion. **The tell:** if the reason a render got rejected lives in a review thread rather than in something you can count, you don't have failure modes yet. You have a folder of bad PNGs. --- ### Step 2 — Make a mode clear four bars before it gets a training run Not every named failure deserves its own training run, and I got this wrong twice before I had a rule I could apply. The bar has four parts. It has to be describable as a behaviour rather than as a defect. "The hand disappears" is a defect; "treat foreground objects inside the mask as immutable" is a behaviour, and a behaviour is something you can put in a dataset. You have to be able to collect data where that behaviour is the dominant variable. Every image in the occlusion set has a hand or an arm or a piece of jewelry interacting with the garment, so the thing you want learned is present in nearly every gradient step. A set where it appears in a fifth of the images trains a fifth of an adapter. It has to conflict with something you're already training, or it doesn't need to be separate. This is the part people skip. **Separation is a cost you pay to avoid interference, so only pay it when interference is what you're seeing.** And it has to be a base model gap rather than a conditioning gap. Style drift looked like an adapter problem for about a week. It wasn't: the model knew perfectly well what a weave is, it wasn't receiving the reference's structure, and Flux Redux fixed it without training anything. Drape and occlusion cleared all four. Kimono-as-bathrobe cleared none of them, since it's the base model's semantic prior and a few thousand images don't move a prior that size. So the reject pile fans out, most of it into places that are not a training run: ```mermaid flowchart LR R[rejects] --> T{name the mode} T -->|base model gap| D[drape set] T -->|base model gap| O[occlusion set] T -->|conditioning gap| X[Flux Redux] T -->|semantic prior| N[out of scope] D --> A[expert A] O --> B[expert B] A --> M[fuse and serve] B --> M ``` The two arrows that never reach a training run are the point of the diagram. Only the branches that survive all four bars cost you a GPU. **The tell:** if you can only state the failure as a defect and not as an instruction you'd give a person, there is nothing to put in a dataset yet. --- ### Step 3 — Argue the case against splitting before you split One general adapter is a better artifact than two specialists in most of the ways that matter once the research is done, and it deserves a fair hearing before I throw it out. One training run to schedule, one file to version, one thing to load, no merge step. The model you evaluated is the model you serve, which stops being true the moment you fuse two adapters, since the fused weights are a third model nobody trained. No coefficient sitting in the middle of your pipeline for somebody to re-derive when the base checkpoint moves. The scaling argument is the strongest one. Under a single adapter a new failure mode is more data in an existing set. Under a split, it's a new artifact plus a new interaction to test against everything already there. Two adapters is one pair, four is six pairs, and there's no reason to believe pairwise coverage is enough. I rejected all of that on one narrow ground: one adapter demonstrably could not hold both behaviours at rank 32, and I had the flat pleats to show for it. If the mixed set had worked I'd have shipped the mixed set. > The split isn't the better design in the abstract. It's what the training runs left us with. **The tell:** if you can't point at a specific artifact the mixed set produced, flat pleats or a hand back off, you're splitting on taste rather than on evidence. --- ### Step 4 — Train the drape expert on images a square crop would ruin Expert A is 5,000 images of sarees, hanfus, kimonos and complex haute couture. Everything in it was chosen because the garment's shape is not a function of the body underneath it. A t-shirt is roughly a tube on a torso. A saree is fabric obeying gravity, and where it ends up depends on how it was wrapped. *One thing worth stating plainly because we confused ourselves with it internally more than once: this 5,000-image set is not the same 5,000-image corpus we used to seed the scene generation side, the one organised by garment layer into denim, silk and leather. Same round number, completely different data, different purpose, different model. If you're reading both, keep them apart.* Filtering was blunt: resolution, sharpness, aspect ratio. Anything soft or oddly cropped went out, because a blurred pleat teaches the model that pleats are blurred. ```bash # train_drape_lora.sh # Expert A: draping physics. Kohya-ss, single L40S (48GB). # Rank stays at 32 for both experts so the two adapters compose predictably later. accelerate launch --num_cpu_threads_per_process 8 \ sdxl_train_network.py \ --pretrained_model_name_or_path "$FLUX_BASE" \ --train_data_dir "./datasets/drape_5k" \ --output_dir "./out/lora_drape" \ --output_name "expert_a_drape" \ --resolution 1024,1024 \ --network_module networks.lora \ --network_dim 32 \ --network_alpha 16 \ --learning_rate 1e-4 \ --lr_scheduler cosine \ --train_batch_size 4 \ --gradient_accumulation_steps 4 \ --gradient_checkpointing \ --optimizer_type AdamW8bit \ --mixed_precision bf16 \ --save_every_n_epochs 1 \ --caption_extension .txt \ --keep_tokens 1 \ --enable_bucket \ --min_bucket_reso 768 \ --max_bucket_reso 1536 ``` Rank 32 with alpha 16 was not an optimisation, it was a decision to stop optimising. Both experts share it, which means when I later merge them I'm combining two updates of the same shape and the same effective scale, and any difference in their influence comes from a coefficient I set rather than from a training artifact I forgot about. Bucketing matters here more than it does for square product shots. A full-length saree image is tall, and cropping it to a square throws away the drape you're trying to teach. **The tell:** if cropping your training images to square costs you nothing, drape probably isn't the variable you're training. --- ### Step 5 — Put the occlusion behaviour in the supervision, not in the hope Expert B is 3,000 images: hands on hips, arms crossed, jewelry sitting over clothing. The model doesn't have to learn what a hand looks like, it already knows. It has to learn that a hand inside the inpainting mask is not a hole to be filled. The size difference between the two sets isn't a budget compromise, it follows from how much variation each concept carries. Drape is a product of several axes at once: garment family, how the fabric was wrapped, its weight and stiffness, and where the folds land for a given stance. Crossing those eats 5,000 images quickly. Occlusion is one relation, foreground in front of fabric, over the small number of poses commercial photography actually uses, so once you have hands on hips from several angles and distances the next thousand images are mostly the same lesson again. Supply matters too. A usable occlusion image needs the interaction clearly visible and the garment boundary clean enough to mask, and the prep script below discards more than I expected on those grounds. *We stopped adding when new images stopped changing anything we could see in the eval renders, which is a soft criterion, and I'd rather say it's soft than dress it up.* The data prep did most of the work. For every image we produced a garment mask and a foreground mask, and the caption named the interaction explicitly so the concept had a token to attach to. ```python # prep_occlusion_set.py """Build the Expert B training set: garment masks with foreground held out. For each source image we run SAM2 twice, once for the garment and once for the occluding body part or accessory, then subtract the second from the first. The adapter is trained against a mask that already excludes the hand, so the target behaviour ("leave this alone") is present in the supervision rather than being something we hope emerges. """ from pathlib import Path import numpy as np from PIL import Image from sam2.build_sam import build_sam2 from sam2.sam2_image_predictor import SAM2ImagePredictor GARMENT_PROMPTS = ["shirt", "tshirt", "top", "saree", "dress"] FOREGROUND_PROMPTS = ["hand", "fingers", "forearm", "necklace", "bracelet"] DILATE_PX = 9 # inside the 5 to 15 px bleeding zone we use at inference def dilate(mask: np.ndarray, radius: int) -> np.ndarray: """Grow a boolean mask by `radius` pixels using a square structuring element.""" from scipy.ndimage import binary_dilation kernel = np.ones((radius * 2 + 1, radius * 2 + 1), dtype=bool) return binary_dilation(mask, structure=kernel) def build_masks(predictor: SAM2ImagePredictor, image: Image.Image): predictor.set_image(np.array(image)) garment, _, _ = predictor.predict(text=GARMENT_PROMPTS, multimask_output=False) foreground, _, _ = predictor.predict(text=FOREGROUND_PROMPTS, multimask_output=False) garment = dilate(garment[0].astype(bool), DILATE_PX) # The foreground is dilated harder. An under-sized hand mask leaks skin into # the region the model is allowed to repaint, and that is exactly the defect # this expert exists to remove. foreground = dilate(foreground[0].astype(bool), DILATE_PX * 2) return garment & ~foreground, foreground def main(src: Path, dst: Path): predictor = SAM2ImagePredictor(build_sam2("sam2_hiera_l.yaml", "sam2_hiera_large.pt")) dst.mkdir(parents=True, exist_ok=True) kept = skipped = 0 for path in sorted(src.glob("*.jpg")): image = Image.open(path).convert("RGB") target, foreground = build_masks(predictor, image) # If the occluder covers almost nothing, the image teaches nothing. if foreground.mean() < 0.005: skipped += 1 continue Image.fromarray((target * 255).astype(np.uint8)).save(dst / f"{path.stem}_mask.png") image.save(dst / path.name) (dst / f"{path.stem}.txt").write_text("hand resting over garment, hand in front of fabric") kept += 1 print(f"kept {kept} / skipped {skipped} (occluder too small)") if __name__ == "__main__": main(Path("raw/occlusion"), Path("datasets/occlusion_3k")) ``` Training config for Expert B is the same as Expert A apart from the data directory and the output name. That's intentional. The only variable I wanted between the two runs was the data. **The tell:** if the prep script is throwing images away because the occluder is too small, that behaviour was never the dominant variable in them, and they were never going to teach it. --- ### What breaks it **The two adapters degrade each other when they're loaded together.** Both passed on their own. Loaded alone against the base, Expert A gives you pleats with volume and shadows that fall the right way, and Expert B keeps hands in front of fabric. Together they interfere, and that turned out to be a much longer story than the one in this post. I'm not going to pretend it was solved in May. It wasn't. **Sheer fabrics still render opaque.** Chiffon and lace need skin tone blended with fabric rather than one replaced by the other, and neither expert was trained to do that. **Extreme poses still fail.** Both datasets skew toward standing and seated subjects, which is what commercial fashion photography mostly is. The dataset that gave each expert its focus is the same dataset that draws its boundary. Where the split did pay off was the eventual eval, on a 500-image set split evenly between Western casual and global traditional. | condition | occlusion accuracy | SSIM, global traditional half | |---|---|---| | base Flux Fill | 70% | 0.58 | | full pipeline | 92% | 0.85 | Those are pipeline numbers rather than adapter numbers, so read them as the ceiling the split contributed to rather than as a measurement of Expert B alone. The occlusion column in particular is the direct descendant of a 3,000-image dataset that exists only because we bothered to name the failure first. --- ### When one adapter per failure mode is the wrong choice The split is a cost. You pay it to avoid interference, and if interference isn't what you're seeing, you're paying for nothing. - **The modes don't conflict.** If adding data for one behaviour doesn't degrade another, keep one set. A mixed adapter is one run, one file, one thing to load, and no coefficient for somebody to re-derive later. - **It's a conditioning gap, not a base model gap.** Style drift looked like an adapter problem for a week. The model already knew what a weave was, it just wasn't receiving the reference's structure, and Flux Redux fixed it with no training at all. - **It's the base model's semantic prior.** Kimono-as-bathrobe is what the base model believes a kimono is. A few thousand images don't move a prior that size, whichever adapter you put them in. - **You can't collect data where the behaviour dominates.** A set where the target behaviour appears in a fifth of the images trains a fifth of an adapter, and you've spent a training run to get it. - **You need one artifact to serve and you can't afford a merge step.** Fused weights are a third model nobody trained. Two adapters is one interaction to test, four is six, and pairwise coverage is an assumption rather than a guarantee. --- ### The shift I spent weeks adding images to a pile called "hard cases" and the pile got worse at both of the things in it. What finally moved it was noticing that two of the defects in there wanted opposite behaviour from the same weights, and that no mixing ratio makes that go away. - **Name the failure before you collect the data.** "Hard cases" is not a category, and a dataset built from it trains an adapter that's mediocre at several things instead of good at one. - **Split for separate eval, not for separate quality.** When one adapter regresses you know which training run to look at and which images to blame. A single mixed finetune gives you a loss curve and a shrug. - **Keep the shapes identical across experts.** Same rank, same alpha, same optimiser. It costs nothing at training time and it's what makes the adapters comparable when you go to combine them. The next thing was combining them into one model we could actually serve, since loading two adapters and switching between them per request was not a serving story I wanted. That's where the interesting part starts, and it did not go the way I expected. --- Source: https://himanshuat.com/blogs/one-lora-per-failure-mode ======================================================================== --- title: "Optimizing Agent Reliability: Debugging Trajectories and Prompt Engineering" description: "An agent fails on the shape of its own output long before it fails on reasoning. Reading trajectories and writing the prompt as a format contract is what turns something that mostly works into something that holds up run after run." date: "April 12, 2024" url: "https://himanshuat.com/blogs/optimizing-langchain-agent-reliability-debugging-prompt-engineering" --- # Optimizing Agent Reliability: Debugging Trajectories and Prompt Engineering The model emits `Action: calculator` with the input on the same line instead of the next one. Your regex expects the newline. `parseAgentOutput` returns `null`, the loop pushes a correction back into history, and one of your five iterations is gone. Nothing in that failure looks like a model problem. It looks like a parser problem, and it is a format problem. I want to build systems that reason, and the path I keep coming back to is Mixture-of-Experts (MoE) architectures, where specialized agents collaborate on a problem. An expert that fails 30% of the time isn't an expert; it's a liability. > For an LLM agent the prompt isn't instructions wrapped around a program. It *is* the program, and every token in it is a line of that program. ### Chapter 0 — What a trajectory is A trajectory is the full sequence of prompts, responses, and observations for a single decision, read in order. Not the input and the final answer. The middle: what the model said, what tool that got parsed into, what came back, and what the model did with what came back. Almost every intermittent agent failure lives in that middle, which is exactly the part an abstraction layer hides. My aversion to opaque, bloated frameworks comes from this. I need to know *exactly* why an agent made a given decision, and when it fails I need to diagnose it without fighting the tooling. The loop this post is about has one branch and one cycle, and both matter: ```mermaid flowchart TD A[LLM call] --> B{parse output} B -->|final answer| C[return] B -->|action| D[run tool] B -->|unparseable| E[format reminder] D --> F[append observation] E --> F F --> A ``` Every arrow into `parse output` is a place the model can deviate, and every arrow back to `LLM call` carries the deviation forward into the next turn. --- ### Step 1 — Strip the agent down to interactions you can read Frameworks like LangChain are convenient for prototyping, but the extra indirection hides performance bottlenecks and makes low-level debugging painful. When every token and every interaction counts, I want direct API calls and minimal abstractions. Here's a stripped-down TypeScript agent loop that keeps the interactions raw: ```typescript // agent.ts interface Tool { name: string; description: string; execute: (input: string) => Promise; } interface AgentState { thought: string; action?: { tool: string; input: string }; observation?: string; history: { role: 'user' | 'assistant' | 'system'; content: string }[]; } class LeanAgent { private tools: Record; private llmClient: any; // Raw OpenAI, Anthropic, or Agno client private systemPrompt: string; constructor(tools: Tool[], systemPrompt: string, llmClient: any) { this.tools = tools.reduce((acc, tool) => ({ ...acc, [tool.name]: tool }), {}); this.llmClient = llmClient; this.systemPrompt = systemPrompt; } private formatToolsForPrompt(): string { return Object.values(this.tools) .map(tool => { let example = ''; // Add specific few-shot examples directly in tool description for clarity if (tool.name === 'search') { example = `Example: search("latest AI research papers")`; } else if (tool.name === 'calculator') { example = `Example: calculator("2+2*3")`; } return `Tool: ${tool.name}(input: string) -> string\nDescription: ${tool.description}\n${example}`; }) .join('\n\n'); } private async callLLM(messages: { role: 'user' | 'assistant' | 'system'; content: string }[]): Promise { // Direct call to LLM API (e.g., OpenAI chat completions) // Add robust retry logic, timeout, and error handling here try { const response = await this.llmClient.chat.completions.create({ model: "gpt-4o", // Or your preferred fast, capable model messages: messages, temperature: 0.1, // Keep it low for reliability max_tokens: 512, }); return response.choices[0].message.content || ''; } catch (error) { console.error("LLM API error:", error); // Implement circuit breaker or specific fallback throw new Error("Failed to get LLM response"); } } public async run(initialQuery: string, maxIterations: number = 5): Promise { let state: AgentState = { thought: '', history: [{ role: 'system', content: this.systemPrompt + '\n\n' + this.formatToolsForPrompt() }], }; state.history.push({ role: 'user', content: initialQuery }); for (let i = 0; i < maxIterations; i++) { console.log(`--- Agent Iteration ${i + 1} ---`); const llmResponse = await this.callLLM(state.history); // This is where custom parsing comes in const parsedOutput = this.parseAgentOutput(llmResponse); if (parsedOutput && 'answer' in parsedOutput) { console.log(`FINAL ANSWER: ${parsedOutput.answer}`); return parsedOutput.answer; } else if (parsedOutput && 'tool' in parsedOutput) { const { tool, input } = parsedOutput; console.log(`Action: ${tool}, Input: ${input}`); const toolExecutor = this.tools[tool]; if (!toolExecutor) { const observation = `Error: Tool '${tool}' not found. Available tools: ${Object.keys(this.tools).join(', ')}`; state.history.push({ role: 'assistant', content: llmResponse }); state.history.push({ role: 'user', content: `Observation: ${observation}` }); } else { try { const observation = await toolExecutor.execute(input); state.history.push({ role: 'assistant', content: llmResponse }); state.history.push({ role: 'user', content: `Observation: ${observation}` }); } catch (error) { const observation = `Tool '${tool}' failed with error: ${error instanceof Error ? error.message : String(error)}`; state.history.push({ role: 'assistant', content: llmResponse }); state.history.push({ role: 'user', content: `Observation: ${observation}` }); } } } else { console.warn("LLM provided malformed or unparseable output. Retrying with explicit instruction."); state.history.push({ role: 'assistant', content: llmResponse }); state.history.push({ role: 'user', content: "Your previous output was malformed. Please ensure you either use 'Action: [tool]\nAction Input: [input]' or 'FINAL ANSWER: [answer]' format. Do not use any other format." }); } } return "Agent failed to find a final answer within the maximum iterations."; } // Custom output parser (implementation below) private parseAgentOutput(llmOutput: string): { tool: string; input: string } | { answer: string } | null { // ... (implementation detailed in section 5) return null; // Placeholder } } ``` Everything is explicit: the LLM call, the tool lookup, the parse, the branch that handles a tool that doesn't exist. Notice that a missing tool is not an exception. It becomes an observation naming the tools that do exist, and the model gets another turn with that information in front of it. That visibility is the entire reason to hand-roll this. **The tell:** if you can't point at the line where the model's raw string becomes a tool call, you don't have an agent you can debug. You have a dependency. --- ### Step 2 — Trace the trajectory, not the output My instinct is to build everything from scratch, but some external tools earn their keep. For seeing an agent's thought process, LangSmith is a good observability layer, if a bit heavy-handed. I wouldn't build the agent *with* it. It's a useful lens for *watching* the agent's state transitions. Integrating it minimally means wrapping `callLLM` and each tool's `execute` to log inputs, outputs, and errors. ```typescript // Minimal LangSmith integration for observability import { Client, RunTree } from "langsmith"; // Assuming LANGCHAIN_API_KEY and LANGCHAIN_TRACING_V2 are set const langsmithClient = new Client(); async function wrapLLMCallWithTrace( llmClient: any, messages: { role: 'user' | 'assistant' | 'system'; content: string }[], parentRun?: RunTree ): Promise<{ output: string; run: RunTree }> { const run = new RunTree({ name: "LLM_Call", run_type: "llm", inputs: { messages }, parent_run: parentRun, // tags: ["agent-trace"] // Add custom tags for filtering }); try { await run.post(); const response = await llmClient.chat.completions.create({ /* ... */ }); const output = response.choices[0].message.content || ''; run.outputs = { output }; await run.end(); return { output, run }; } catch (error) { run.error = String(error); await run.end(); throw error; } } // In the Agent's run method: // let currentRun: RunTree | undefined; // Pass this through iterations // ... // const { output: llmResponse, run: llmRun } = await wrapLLMCallWithTrace(this.llmClient, state.history, currentRun); // currentRun = llmRun; // Update current run for next step ``` The `parentRun` threading is the part that matters. Without it you get a pile of unrelated LLM calls; with it you get one tree per query, which is what makes a trajectory readable rather than a log. Five things a trace surfaces that a log of inputs and answers cannot: * **Deviation from instructions.** The model outputs a thought-action sequence different from the one the prompt asked for. * **Tool misuse.** The agent calls the wrong tool, or passes malformed arguments to the right one. This usually points to ambiguous tool descriptions or too few examples. * **Observation misinterpretation.** The agent receives an observation and then acts as if it saw something else. * **Infinite loops.** The agent repeats the same actions without making progress. * **Hallucinated tools.** The model invents a tool that doesn't exist. Each is a distinct fix. Tool misuse is a description problem. Observation misinterpretation is a prompt problem. Telling them apart requires seeing the observation the model actually received, in position, between the two turns. **The tell:** if your logs record inputs and final answers but not the observation in between, you can't distinguish tool misuse from observation misreading. They look identical from the outside. --- ### Step 3 — Write the system prompt as a format contract Once a trace highlights a failure, the fix almost always lives in the prompt. The system prompt sets identity and goals, and, critically, **output format constraints**. ```typescript const systemPrompt = ` You are a highly efficient and accurate problem-solving agent. Your goal is to precisely answer the user's query by breaking it down into steps, using available tools, and providing a concise FINAL ANSWER. Strictly follow this thought-action-observation loop: 1. **Thought**: You must always first reflect on the current state, what you need to do next, and which tool to use. 2. **Action**: If you need to use a tool, output 'Action: [tool_name]\nAction Input: [tool_input]'. 3. **Observation**: This will be provided by the system after your action. 4. **FINAL ANSWER**: Once you have fully solved the user's request and have the complete answer, output 'FINAL ANSWER: [your_answer]'. Do NOT elaborate or provide conversational text outside of your 'Thought' or 'FINAL ANSWER'. You MUST ONLY use the tools provided. DO NOT hallucinate tool names or arguments. If you get stuck or receive an error, try to recover or state a FINAL ANSWER if appropriate. `; ``` Four things carry the weight: * **Explicit format enforcement.** The two legal output shapes are named as literals, not described. Anything else is a violation the parser has to catch. * **Role and goal clarity.** "You are a highly efficient and accurate problem-solving agent." * **Constraint repetition.** Repeat the critical constraints. Models forget instructions mid-response, and the format constraint is the one that costs you an iteration when it's forgotten. * **Temperature.** Keep it low, `0.1` or `0`, trading creativity for reliability. **The tell:** if the format constraint appears once, at the top of a long system prompt, count the tokens between it and the model's turn. That distance is the deviation rate you're paying for. --- ### Step 4 — Put the invocation example inside the tool description Ambiguous tool descriptions are a prime source of agent failure. Each tool needs a clear purpose, a clear argument structure, and **few-shot examples embedded directly in the prompt**. ```typescript // Inside formatToolsForPrompt() method or similar utility const tools: Tool[] = [ { name: 'search', description: 'Searches the web for factual information. Use this when you need current data or to verify facts.', execute: async (query: string) => `Simulated web search for "${query}" resulted in "LLMs are complex neural networks."`, }, { name: 'calculator', description: 'Evaluates mathematical expressions. Input must be a valid arithmetic string.', execute: async (expression: string) => `Simulated calculation of "${expression}" resulted in "10"`, }, // ... more tools ]; // Example of how they render in the prompt: // Tool: search(input: string) -> string // Description: Searches the web for factual information. Use this when you need current data or to verify facts. // Example: search("latest AI research papers") // // Tool: calculator(input: string) -> string // Description: Evaluates mathematical expressions. Input must be a valid arithmetic string. // Example: calculator("2+2*3") ``` The `Example:` line does a lot of work. A description tells the model when to reach for a tool. An example tells it what the argument looks like when it does, and that cuts ambiguity far better than description alone. Note where the example lives: in `formatToolsForPrompt`, next to the signature, not in a separate block of demonstrations further up the prompt. The model reads the tool and its invocation pattern in the same breath. **The tell:** if a tool's description explains what it does but never shows one concrete call, the model is guessing at the argument shape, and the guess will be plausible enough to survive your parser. --- ### Step 5 — Parse defensively and feed the failure back The LLM is a language model, not a JSON serializer. It will deviate from your output format, and a single missing newline or extra word can crash the agent. You need a parser built for your agent's exact output, not a generic framework one. ```typescript // Inside the LeanAgent class private parseAgentOutput(llmOutput: string): { tool: string; input: string } | { answer: string } | null { llmOutput = llmOutput.trim(); // Always trim whitespace // Attempt to parse Final Answer first const finalAnswerMatch = llmOutput.match(/^FINAL ANSWER:\s*(.*)/is); if (finalAnswerMatch) { return { answer: finalAnswerMatch[1].trim() }; } // Attempt to parse Action // Regex is designed to be robust against slight variations in newlines/spacing const actionMatch = llmOutput.match(/Action:\s*([a-zA-Z0-9_]+)\s*\nAction Input:\s*(.*)/is); if (actionMatch) { return { tool: actionMatch[1].trim(), input: actionMatch[2].trim(), }; } // Fallback: If neither matches, try to salvage console.warn("LLM output malformed, attempting heuristic recovery:", llmOutput); // Heuristic 1: If it starts with "Thought:" and then has "Action:" const thoughtActionMatch = llmOutput.match(/Thought:\s*.*?\s*Action:\s*([a-zA-Z0-9_]+)\s*\nAction Input:\s*(.*)/is); if (thoughtActionMatch) { console.warn("Recovered from Thought-Action format."); return { tool: thoughtActionMatch[1].trim(), input: thoughtActionMatch[2].trim(), }; } // Heuristic 2: Simple case where Action/Input might be on same line or separated differently const simpleActionMatch = llmOutput.match(/Action:\s*([a-zA-Z0-9_]+)\s*Input:\s*(.*)/is); if (simpleActionMatch) { console.warn("Recovered from simple Action-Input format."); return { tool: simpleActionMatch[1].trim(), input: simpleActionMatch[2].trim(), }; } // If all recovery attempts fail console.error("Failed to parse LLM output after all attempts:", llmOutput); return null; // Signal a parsing failure, which the agent loop should handle } ``` The ordering is deliberate. `FINAL ANSWER` is checked first and anchored to the start of the string, so a model that narrates its way toward an action can't accidentally terminate the run. The strict `Action` pattern comes next. Only then do the two heuristics run, each behind its own `console.warn`, so a recovery is always visible in the trace rather than silently absorbed. When everything fails, the parser returns `null` and the loop takes over: log the malformed output, push the model's own response back into history, and follow it with a reminder of the two legal formats. That self-correction is what keeps a single bad response from ending the run. **The tell:** if a malformed response can end a run, your parser is part of the model's failure surface, not a defense against it. --- ### Four things to add this week These are all in the code above, and each is a few lines. 1. **Trim before you match.** Leading whitespace defeats an anchored `^FINAL ANSWER` regex and turns a correct response into a parse failure. 2. **Name the available tools in the not-found observation.** `Available tools: ${Object.keys(this.tools).join(', ')}` turns a hallucinated tool from a dead end into a correction the model can act on. 3. **Wrap `execute`, not just `callLLM`.** A trace with the LLM calls but not the tool results shows you what the model said and hides what it saw. 4. **Drop the temperature.** `temperature: 0.1` on a format-constrained agent costs you nothing you want. --- ### What breaks it **Heuristic recovery hides the defect it recovers from.** Every successful salvage in `parseAgentOutput` is evidence that the prompt's format contract is not holding. The run completes, so nothing forces you to look. The `console.warn` on each heuristic exists so that the recovery rate stays visible; if you never read those warnings, you've built a parser that quietly absorbs a prompt regression. **The iteration cap returns a string, not an error.** When `maxIterations` runs out, `run` returns `"Agent failed to find a final answer within the maximum iterations."` — a `Promise`, the same type as a real answer. A caller that doesn't check for that exact sentence will treat a failed run as a successful one. **Every retry makes the next turn harder.** The malformed-output branch pushes the bad response into history and then the correction after it. The model's own deviation is now context it will read on the next turn. A run that deviates twice is carrying two examples of the wrong format alongside your instruction to use the right one. **The LLM error path drops the detail.** `callLLM` logs the caught error and then throws a fresh `Error("Failed to get LLM response")`. Rate limit, timeout, and malformed request all arrive at the caller looking identical. The comment in that block names the fix, a circuit breaker or a specific fallback, and it's a comment rather than code for a reason. --- ### When a hand-rolled lean agent is the wrong choice The argument above is for control, and control has a maintenance bill attached. Sometimes the bill is larger than the problem. - **You're still exploring what the agent should do.** Frameworks are convenient for prototyping precisely because you aren't yet paying attention to the interaction points. Get the shape right first, then strip it down. - **The output format keeps moving.** Every format change is your regex to rewrite, plus two heuristics to re-derive. A parser you own is only cheaper than a generic one while the format is stable. - **There's one tool and one call.** No branch, no cycle, nothing to trace. The loop machinery, the parse ladder, and the trace tree are overhead when the agent is a function call with a system prompt. - **The failure is reasoning, not format.** If the model picks the wrong tool for reasons a correct description and a worked example don't fix, no amount of defensive parsing helps. A parser can only catch a badly shaped answer, never a badly chosen one. --- ### The shift The instinct when an agent misbehaves is to reach for a better model. Most of the failures in this post survive a model upgrade, because they're failures of format, description, and visibility rather than capability. What changes tomorrow: - **Read trajectories, not outputs.** The observation between two turns is where the diagnosis lives; input and answer alone can't tell tool misuse from misreading. - **Treat the prompt as a performance-critical function.** Format constraints stated as literals, repeated, close to the model's turn. - **Assume malformed output and design for it.** Parse defensively, log every recovery, feed the correction back into the loop. If the goal is to assemble many specialized experts into something larger, each one has to be a dependable module first. > An agent that mostly works is a component you can't build on. Reliability is the format contract, enforced at every turn. --- Source: https://himanshuat.com/blogs/optimizing-langchain-agent-reliability-debugging-prompt-engineering ======================================================================== --- title: "Precision Tuning: Optimizing the Retriever vs. the Generator in Your RAG Pipeline" description: "The wrong chunk comes back ranked first and the generator answers on top of it anyway. Fixing that means deciding which half of the pipeline is actually broken, the retriever or the generator, and then paying for the right one in the right order." date: "November 06, 2025" url: "https://himanshuat.com/blogs/optimizing-rag-fine-tuning-retriever-vs-generator" --- # Precision Tuning: Optimizing the Retriever vs. the Generator in Your RAG Pipeline Two terms are distinct in your corpus. To an embedding model trained on Wikipedia and web crawl, they read as synonyms. So the wrong chunk comes back ranked first, the generator answers on top of it, and nothing in the pipeline errors. You get a fluent paragraph that is confidently about the wrong document. Off-the-shelf RAG is a starting point, not a destination. Spin up an embedding model, hook it to a vector store, pair it with a general-purpose LLM, and you get okay results. For general questions, okay is fine. > Vanilla RAG doesn't fail because the models are weak. It fails because they're general, and generality is a property you train out of one component at a time. ### Chapter 0 — The two stages you're allowed to touch A RAG pipeline is two stages, and that split is the whole reason this is tractable. **Retrieval** fetches relevant chunks from a knowledge base. Usually a **bi-encoder** (a `SentenceTransformer`, say) embeds queries and documents into a shared vector space, and cosine similarity in that space stands in for relevance. **Generation** is a **generative LLM** synthesizing an answer from the retrieved context plus the query. Each stage fails differently, costs differently, and takes different training data. Off-the-shelf embedding models handle general semantic similarity well and miss the jargon and specific relationships of a niche domain: medical diagnoses, legal precedents, obscure physics papers. General-purpose LLMs write fluent text and hallucinate domain-inappropriate facts, miss your required answer format, or fall back on generic responses. They haven't learned the voice of your domain. The job is to turn a generalist into a specialist, component by component. --- ### Step 1 — Fix the retriever before you touch the generator The retriever is the foundation. If it fetches the wrong context, even a great LLM will hallucinate or answer generically, because it's synthesizing faithfully from the wrong material. Improving retrieval improves the *quality of information* the generator has to work with, which means it improves every step downstream of it. Fine-tuning an embedding model is also cheaper and faster than fine-tuning a large LLM. The cheap fix is also the one with the widest blast radius, which almost never happens and is worth exploiting when it does. **The tell:** paste the retrieved chunks next to the bad answer. If the answer isn't in them, the generator was never your problem. --- ### Step 2 — Generate the retrieval training data instead of labeling it Fine-tuning a bi-encoder means teaching it *your domain's* notion of relevance. The hard part isn't the training loop. It's the pairs. Hand-labeling query-document relevance across a large corpus is too slow and too expensive to finish. So invert it: take each document you already have and use a general-purpose LLM (GPT-4, Llama 3) to generate the plausible queries that should retrieve it. ```python import openai # Or your preferred LLM API client, e.g., vLLM for local models import json from typing import List, Dict # Assume 'corpus' is a list of dictionaries, each with 'id' and 'text' corpus: List[Dict[str, str]] = [ {"id": "doc1", "text": "The Agno framework prioritizes developer control and minimal abstraction..."}, {"id": "doc2", "text": "Attention mechanisms revolutionized sequence modeling, allowing models to weigh different parts of the input..."}, # ... more documents ] def generate_synthetic_queries(document_text: str, num_queries: int = 3) -> List[str]: """Generates synthetic queries for a given document using an LLM.""" prompt = f""" You are an expert search query generator. Given the following document, generate {num_queries} diverse and highly relevant search queries that a user might ask to find this specific document. Focus on key entities, concepts, and relationships within the document. Each query should be short and specific. Document: --- {document_text} --- Generate {num_queries} queries, each on a new line, prefixed with '- ': """ try: response = openai.chat.completions.create( model="gpt-4o", # Or "llama3", etc. messages=[ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": prompt} ], temperature=0.7, max_tokens=200 ) queries = [line.strip('- ').strip() for line in response.choices[0].message.content.split('\n') if line.strip().startswith('- ')] return queries[:num_queries] # Ensure we return exactly num_queries except Exception as e: print(f"Error generating queries: {e}") return [] # Example usage to create (query, positive_document) pairs training_data = [] for doc in corpus: queries = generate_synthetic_queries(doc['text']) for query in queries: training_data.append({"query": query, "positive_document": doc['text']}) print(f"Generated {len(training_data)} synthetic (query, positive_document) pairs.") ``` Positives alone don't teach a metric space anything. You need *negatives*: documents that are irrelevant to a given query, either random or hard. Hard negatives are the semantically similar ones that are still wrong, and they're what pushes the model to tell fine-grained relevance apart. You find them by retrieving the top-k for a query and dropping the known positives. Then train contrastively. Pull the query and its positive document together in vector space, push the negatives away. Triplet loss is a common choice. ```python import torch from torch.utils.data import Dataset, DataLoader from sentence_transformers import SentenceTransformer, InputExample, losses from sentence_transformers.evaluation import InformationRetrievalEvaluator from transformers import AutoTokenizer, AutoModel # For manual control if not using sentence_transformers # Assuming 'training_data' from above, now with 'query', 'positive_document', 'negative_document' # Example: training_data = [{"query": "...", "positive": "...", "negative": "..."}, ...] class TripletDataset(Dataset): def __init__(self, data): self.data = data def __len__(self): return len(self.data) def __getitem__(self, idx): item = self.data[idx] return InputExample(texts=[item['query'], item['positive_document'], item['negative_document']]) # Initialize a base bi-encoder model (e.g., from Hugging Face) model_name = 'sentence-transformers/all-MiniLM-L6-v2' # Start with a good base model = SentenceTransformer(model_name) # Prepare data for DataLoader train_examples = [ InputExample(texts=[d['query'], d['positive_document']], label=1.0) # Simplified for contrastive loss for d in training_data ] # For a full triplet loss, you'd need the negative pairs explicitly in InputExample. # A more robust approach directly uses the MultipleNegativeRankingLoss train_dataloader = DataLoader(train_examples, shuffle=True, batch_size=32) # Define the loss function train_loss = losses.MultipleNegativeRankingLoss(model=model) # Fine-tune the model num_epochs = 5 warmup_steps = int(len(train_dataloader) * num_epochs * 0.1) # 10% of total training steps model.fit( train_objectives=[(train_dataloader, train_loss)], epochs=num_epochs, warmup_steps=warmup_steps, output_path='./fine_tuned_retriever', show_progress_bar=True ) print("Retriever fine-tuning complete. Model saved to ./fine_tuned_retriever") ``` `sentence-transformers` is convenient and it hides a lot. For full control you'd load `AutoTokenizer` and `AutoModel` directly, tokenize yourself, compute embeddings with `model(**inputs).last_hidden_state[:, 0]` (CLS token or mean pooling), and write the contrastive loss in PyTorch. That's the version where you can see what the objective actually is. Fine-tune the bi-encoder when retrieval accuracy is low and your domain vocabulary diverges from general language. Those two conditions travel together, and neither one alone is enough. **The tell:** if you're still hand-labeling relevance pairs, you'll run out of patience long before you run out of corpus. Synthetic generation is what makes the dataset finishable. --- ### Step 3 — Add a cross-encoder only when precision is the bottleneck Even with a fine-tuned bi-encoder, the top-k can still hold noise. That's structural, not a tuning failure. A bi-encoder trades interaction depth for speed: it encodes query and document separately, so they never see each other. A **cross-encoder** takes the query and document *together* and scores them in one forward pass, letting them interact at the token level. More compute, more precision. It sits behind the bi-encoder rather than replacing it. Fetch 50-100 candidates with the (fine-tuned) bi-encoder, score each query-document pair with the cross-encoder on a 0-to-1 scale, and pass the top 5-10 to the generative LLM. Data generation mirrors the bi-encoder's, except now you need (query, document, relevance_score) triplets, which means using the LLM as an annotator rather than a query writer. ```python # Re-using the generate_synthetic_queries function, but now focusing on scoring existing pairs. # Instead of just generating queries for a doc, we can generate a score for a (query, doc) pair. def generate_cross_encoder_data(query: str, document_text: str) -> Dict[str, any]: """ Generates a relevance score and potentially alternative queries for a (query, document) pair. This uses an LLM to act as a human annotator. """ prompt = f""" Given the following search query and document, evaluate their relevance. Assign a relevance score from 0 (completely irrelevant) to 1 (perfectly relevant). Also, identify why it's relevant/irrelevant and suggest a more precise query if the current one is vague. Query: "{query}" Document: --- {document_text} --- Output a JSON object with 'relevance_score' (float), 'explanation' (string), 'suggested_query' (string, or null if perfect): """ try: response = openai.chat.completions.create( model="gpt-4o", messages=[ {"role": "system", "content": "You are a helpful assistant that analyzes text relevance."}, {"role": "user", "content": prompt} ], temperature=0.3, response_format={"type": "json_object"} ) result = json.loads(response.choices[0].message.content) # Basic validation if 'relevance_score' in result and isinstance(result['relevance_score'], (int, float)): return {"query": query, "document": document_text, "score": float(result['relevance_score'])} else: print(f"Invalid JSON format for relevance score: {result}") return None except Exception as e: print(f"Error generating cross-encoder data: {e}") return None # Example loop (this would be expensive, usually done with a mix of synthetic and actual feedback) cross_encoder_training_data = [] # For each (query, positive_doc) pair, add it with score ~1.0 # For each (query, negative_doc) pair, add it with score ~0.0 (or retrieve weak negatives and score them) # For demonstration, let's take some pairs and score them. sample_pair = {"query": "optimization techniques RAG pipeline", "document": corpus[0]['text']} # Assuming corpus[0] is relevant scored_data = generate_cross_encoder_data(sample_pair["query"], sample_pair["document"]) if scored_data: cross_encoder_training_data.append(scored_data) # Training (simplified, using sentence-transformers for convenience, but can be done with raw transformers) from sentence_transformers import CrossEncoder cross_encoder_model_name = 'cross-encoder/ms-marco-MiniLM-L-6-v2' # Good starting point cross_encoder = CrossEncoder(cross_encoder_model_name, num_labels=1) # num_labels=1 for regression (score) # For training, you need (query, document, score) triplets. # The data would look like: [InputExample(texts=['query', 'document'], label=score), ...] # This requires preparing a dataset of such InputExamples. # Then: cross_encoder.fit(train_dataloader, ...) ``` The re-ranker earns its latency in exactly one situation: irrelevant context is costly, as in legal or medical work where accuracy is critical, and your bi-encoder already returns *most* of the relevant documents so the remaining job is surfacing the *best* ones. **The tell:** a re-ranker can only reorder what it was handed. If the right document never made the top 50-100, that's a recall problem, and recall is the bi-encoder's job. --- ### Step 4 — Teach the generator to stay inside the context Once retrieval is good, the generator is what stands between a correct chunk and a correct answer. If the context is relevant and the model still answers generically or invents things, this is where the work is. Fine-tuning a large LLM is expensive. LoRA and QLoRA make it accessible by training only small adapter layers instead of the full weights. What you're teaching, specifically: - Use the provided context rather than the parametric memory it arrived with. - Match your domain's output format and style. - Say the context isn't enough, instead of inventing an answer. The data problem returns in a harder form. You need (query, context, ideal_answer) triplets where the ideal answer sticks strictly to the provided context, which is a constraint you have to put in the generating prompt because nothing else enforces it. ```python def generate_instruction_tuning_data(query: str, context: str) -> Dict[str, str]: """Generates an ideal answer based *only* on the provided context.""" prompt = f""" You are an expert assistant for a highly specialized domain. Your task is to answer user queries truthfully and concisely, strictly based on the provided context. If the answer cannot be found within the context, state "Information not available in the provided context." Do not introduce any outside information. Maintain a formal and precise tone. Query: "{query}" Context: --- {context} --- Answer: """ try: response = openai.chat.completions.create( model="gpt-4o", # Use the best available model for data generation messages=[ {"role": "system", "content": "You are a helpful assistant that generates high-quality instruction tuning data."}, {"role": "user", "content": prompt} ], temperature=0.7, max_tokens=500 ) generated_answer = response.choices[0].message.content.strip() # Optionally, add a check to ensure the generated_answer is not trivial like "I cannot answer." # Or you can post-filter. return {"query": query, "context": context, "answer": generated_answer} except Exception as e: print(f"Error generating instruction tuning data: {e}") return None # Example: Generate data by pairing synthetic queries with their relevant docs (from retriever fine-tuning) instruction_data = [] for item in training_data: # Assuming 'training_data' has 'query' and 'positive_document' # Use the positive_document as context for the generator data_point = generate_instruction_tuning_data(item['query'], item['positive_document']) if data_point: instruction_data.append(data_point) print(f"Generated {len(instruction_data)} instruction tuning data points.") ``` The training side is more involved than the retriever's: load the base model, quantize it for QLoRA, configure the adapters, run the loop. ```python from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, Trainer from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training import torch from datasets import Dataset # Hugging Face datasets library # Assume 'instruction_data' is a list of {"query": ..., "context": ..., "answer": ...} dicts # Format for instruction tuning: combine into a single text string def format_instruction(item): prompt_template = ( f"### Instruction:\n{item['query']}\n\n" f"### Context:\n{item['context']}\n\n" f"### Answer:\n{item['answer']}" ) return {"text": prompt_template} formatted_dataset = Dataset.from_list([format_instruction(d) for d in instruction_data]) # Choose your base LLM (e.g., Llama-2, Mistral, Gemma) model_id = "meta-llama/Llama-2-7b-hf" # Requires Hugging Face token for Llama # model_id = "mistralai/Mistral-7B-v0.1" # Load base model and tokenizer tokenizer = AutoTokenizer.from_pretrained(model_id) tokenizer.pad_token = tokenizer.eos_token # Important for some models # Load model in 4-bit precision for QLoRA model = AutoModelForCausalLM.from_pretrained( model_id, load_in_4bit=True, torch_dtype=torch.bfloat16, device_map="auto" ) model.config.use_cache = False # Disable cache for gradient checkpointing # Prepare model for k-bit training model = prepare_model_for_kbit_training(model) # Configure LoRA lora_config = LoraConfig( r=8, # LoRA attention dimension lora_alpha=16, # Alpha parameter for LoRA scaling target_modules=["q_proj", "k_proj", "v_proj", "o_proj"], # Apply LoRA to attention weights lora_dropout=0.05, bias="none", task_type="CAUSAL_LM", ) model = get_peft_model(model, lora_config) # Tokenize the dataset def tokenize_function(examples): return tokenizer(examples["text"], truncation=True, max_length=1024, padding="max_length") tokenized_dataset = formatted_dataset.map(tokenize_function, batched=True, remove_columns=["text"]) # Training arguments training_args = TrainingArguments( output_dir="./fine_tuned_generator", num_train_epochs=3, per_device_train_batch_size=4, gradient_accumulation_steps=2, learning_rate=2e-4, logging_steps=10, save_steps=100, fp16=True, # Use mixed precision optim="paged_adamw_8bit", ) # Trainer trainer = Trainer( model=model, args=training_args, train_dataset=tokenized_dataset, tokenizer=tokenizer, ) trainer.train() trainer.save_model("./fine_tuned_generator_lora_adapters") print("Generator fine-tuning (LoRA) complete. Adapters saved.") ``` Reach for this when retrieval is already solid but output quality is not: you need a specific voice, complex formatting, or a model that says "Information not available" when the context is thin. **The tell:** if the answer is wrong in a way the retrieved context contradicts, that's the generator. If it's wrong in a way the retrieved context also gets wrong, you're back in Step 1. --- ### Step 5 — Sequence the two, and measure in between This was never either/or. The question is where the next unit of effort goes, and measurement decides that, not preference. Start with the retriever, always. Then measure it, and let the number pick the next move. ```mermaid flowchart TD A[measure retriever] --> B{recall ok?} B -->|no| C[fine-tune bi-encoder] C --> A B -->|yes| D{top-k precise?} D -->|no| E[add cross-encoder] E --> A D -->|yes| F[measure generator] F --> G{grounded answers?} G -->|no| H[instruction-tune LLM] H --> F G -->|yes| I[ship] ``` The loop matters more than the boxes. Every fine-tune sends you back to the measurement that justified it, because a component you tuned is a component whose failure signature just changed. Each symptom maps to exactly one component, which is the payoff for keeping the stages separate: | symptom | component | fix | |---|---|---| | Low recall or precision on retrieval | bi-encoder | synthetic queries, contrastive fine-tune | | Good recall, noisy top-k | re-ranker | cross-encoder, if latency allows | | Context relevant, answer generic or hallucinated | generator | instruction tuning | | Answer correct, format wrong | generator | instruction tuning | For state-of-the-art, do all of it, in stages: bi-encoder first to get retrieval solid, cross-encoder second and optionally, generator last for context adherence, style, and domain reasoning. The order isn't arbitrary. It's the cost curve: | stage | cost | what it buys | ROI | |---|---|---|---| | Bi-encoder fine-tune | low | large effect on overall quality | high | | Cross-encoder fine-tune | more than the bi-encoder, adds latency | sharper top-k precision | medium | | Generator fine-tune | high — data generation plus compute | voice, format, context adherence | high, but only once the retriever is good | Generator fine-tuning is the expensive one, and it's the one whose value depends entirely on a prerequisite. Doing it first means paying the most for the least. **The tell:** if you can't say which component produced the last bad answer, you aren't ready to fine-tune either of them. --- ### What breaks it **The data generator is the ceiling.** Every training set here is written by a general-purpose LLM: the queries in Step 2, the relevance scores in Step 3, the ideal answers in Step 4. The retriever learns that model's notion of relevance, and the generator learns that model's idea of a good answer. If the annotator doesn't understand your domain, neither will anything you train on its output. **Hard negatives are only as good as your labels.** The procedure is to retrieve the top-k and drop the known positives. Anything relevant but unlabeled stays in the pile and gets trained against as a negative. You are teaching the model that a correct document is wrong. **The convenience layer hides the objective.** The bi-encoder snippet above builds `InputExample`s from query and positive alone, marked "simplified for contrastive loss", and hands them to `MultipleNegativeRankingLoss`, while the `TripletDataset` next to it expects an explicit negative. Those are two different training objectives with the same-looking call site. `sentence-transformers` will run either one without complaint, which is exactly why the raw `AutoModel` version is worth writing once. --- ### When precision tuning is the wrong choice Most pipelines don't need any of this. Fine-tuning components you didn't need to tune costs money and adds surfaces that can drift. - **Your questions are general.** Off-the-shelf embeddings handle general semantic similarity well. If your users ask general questions, okay results are the correct results, and this whole post is overhead. - **Your domain vocabulary doesn't diverge from general language.** Bi-encoder fine-tuning buys you a domain-specific notion of relevance. If your domain's notion is the general one, you're paying to relearn what the base model already knows. - **The document isn't in the corpus.** No amount of tuning retrieves something that was never indexed. Retrieval failures and coverage failures look identical from the answer side and have nothing in common upstream. - **You can't spend a second forward pass.** The cross-encoder is a latency purchase. If the budget isn't there, tune the bi-encoder and stop; a re-ranker you can't afford to run isn't a design option. - **Retrieval is good and the answers are already correct in content and format.** Generator fine-tuning is the highest-cost stage on the list. Absent a specific complaint about voice, format, or groundedness, there's nothing for it to fix. --- ### The shift Optimizing RAG keeps coming back to modular specialization. The pipeline works best when each component is tuned for its own role in a specific domain, rather than left general and asked to cope. Which means the work is diagnostic before it's computational. Find out which of the two components is lying to you, then fix only that one. - **Profile before you tune.** Retrieval, re-ranking, or generation — name the bottleneck first. - **Generate the data for that bottleneck, not for the pipeline.** Queries for the bi-encoder, scores for the re-ranker, grounded answers for the generator. - **Apply one strategy to one component, and no more.** Then re-measure, because you just changed what the failures look like. The frontier I find worth working on is specialization: models matched precisely to their purpose. --- Source: https://himanshuat.com/blogs/optimizing-rag-fine-tuning-retriever-vs-generator ======================================================================== --- title: "Production-Grade Agent Architecture: Implementing Long-Term Memory and Human-in-the-Loop" description: "A stateless agent forgets every decision the moment the process dies, and an agent holding a deploy tool will ship to production without asking anyone. Two pieces of engineering close both gaps: semantic memory over a vector store, and an approval gate that lives in the graph's state instead of in a callback." date: "April 19, 2024" url: "https://himanshuat.com/blogs/production-grade-agent-architecture-memory-human-in-the-loop" --- # Production-Grade Agent Architecture: Implementing Long-Term Memory and Human-in-the-Loop The agent picks `deploy_code`, fills in a project id and a version, and calls it. Nothing between that decision and production asks a human anything. Give the same loop an email tool or a payments tool and it behaves identically, because from inside the loop those calls look the same as reading a row from a database. Restart the process afterwards and it has forgotten the whole thing. > Getting from prototype to production isn't a bigger model or a faster GPU. It's two organs the prototype never needed: memory that outlives the process, and a gate that outranks the model. ### Chapter 0 — The two organs Without persistent memory the agent is an amnesiac. Without human oversight it's a loose cannon. Those are different failures and they need different machinery. **Long-term memory** is the `hippocampus` equivalent. Not a key-value store — semantic retrieval, where the agent queries a concept and gets back related experiences even when the exact keywords never appear. That mirrors how associations work in a brain, which is where I came into this from: what I want to build is a scaled-up specialised cortex driven by a Mixture-of-Experts (MoE) architecture. **Human-in-the-loop** is the `prefrontal cortex` analogue: executive oversight. Autonomy is the goal, but an agent that can deploy code, send critical emails, or move money without supervision is reckless. HITL is the checkpoint where a proposed action gets measured against human intent. In the MoE agent, each expert learns from its own past interactions, and a dedicated Safety Expert owns the human interaction and the pause. Everything below is raw APIs and TypeScript. High-level frameworks trade away control and performance, and both of those are the point here. --- ### Step 1 — Make memory semantic, not keyed Three operations, in order: 1. **Embed.** Turn textual experiences — thoughts, observations, conversations — into dense vectors with an embedding model. The vector carries the meaning. 2. **Store.** Persist the vectors in a database built for similarity search. 3. **Retrieve.** Embed the query, search for the nearest vectors, return the records behind them. Define the boundary before the implementation, because the embedding model and the vector store are both things you will swap. ```typescript // interfaces/memory.ts export interface MemoryRecord { id: string; content: string; timestamp: number; metadata?: Record; embedding?: number[]; // Stored for retrieval if necessary, but primarily for DB } export interface IVectorStoreClient { /** * Stores a batch of memory records after embedding their content. * @param records Records to store. * @returns Promise resolving to void. */ add(records: Omit[]): Promise; /** * Retrieves the top_k most semantically similar records to the query. * @param query The text query to search for. * @param topK The number of similar records to retrieve. * @returns Promise resolving to an array of MemoryRecord. */ search(query: string, topK: number): Promise; } ``` Two methods. `add` takes content and handles embedding internally, so callers never hold a vector. `search` takes text and a `topK`, never a vector either. **The tell:** if the only way to find a memory is the key you filed it under, you built a cache. Semantic memory answers queries whose words are nowhere in the stored record. --- ### Step 2 — Call the embedding and vector APIs directly The embedding model and the vector store drive performance, so the abstraction layers that sit between you and them are latency you cannot see and cannot tune. Skip them. Assume a `Cohere` or `OpenAI` embeddings API and a vector database like `Qdrant`, or a local `Faiss` index when you need high throughput at low latency. ```typescript // services/embeddingService.ts import axios from 'axios'; class EmbeddingService { private readonly apiKey: string; private readonly apiUrl: string; constructor(apiKey: string, apiUrl: string = "https://api.openai.com/v1/embeddings") { this.apiKey = apiKey; this.apiUrl = apiUrl; } /** * Generates embeddings for a batch of texts. * Batched because one round trip per memory is the cost that kills you. */ public async embedBatch(texts: string[]): Promise { if (texts.length === 0) return []; try { const response = await axios.post( this.apiUrl, { model: "text-embedding-3-small", // Or other performant model input: texts, }, { headers: { 'Authorization': `Bearer ${this.apiKey}`, 'Content-Type': 'application/json', }, } ); return response.data.data.map((item: any) => item.embedding); } catch (error) { console.error("Error generating embeddings:", error); throw new Error("Failed to generate embeddings."); } } } // services/vectorStoreClient.ts import { v4 as uuidv4 } from 'uuid'; import { IVectorStoreClient, MemoryRecord } from '../interfaces/memory'; import { EmbeddingService } from './embeddingService'; import { QdrantClient } from '@qdrant/js-client-rest'; // Example: Using Qdrant export class QdrantVectorStoreClient implements IVectorStoreClient { private readonly embeddingService: EmbeddingService; private readonly qdrantClient: QdrantClient; private readonly collectionName: string; constructor(embeddingService: EmbeddingService, qdrantHost: string, collectionName: string = "agent_memories") { this.embeddingService = embeddingService; this.qdrantClient = new QdrantClient({ host: qdrantHost }); this.collectionName = collectionName; this.initializeCollection().catch(console.error); // Ensure collection exists } private async initializeCollection() { const collectionExists = (await this.qdrantClient.getCollections()).collections.some( (c) => c.name === this.collectionName ); if (!collectionExists) { await this.qdrantClient.createCollection(this.collectionName, { vectors_config: { size: 1536, distance: 'Cosine' }, // Match embedding model dimension }); console.log(`Qdrant collection '${this.collectionName}' created.`); } } public async add(records: Omit[]): Promise { if (records.length === 0) return; const contents = records.map(r => r.content); const embeddings = await this.embeddingService.embedBatch(contents); const points = records.map((record, index) => ({ id: uuidv4(), vector: embeddings[index], payload: { content: record.content, timestamp: record.timestamp, metadata: record.metadata, }, })); await this.qdrantClient.upsert(this.collectionName, { wait: true, batch: { ids: points.map(p => p.id), vectors: points.map(p => p.vector), payloads: points.map(p => p.payload), }, }); } public async search(query: string, topK: number): Promise { if (!query) return []; const [queryEmbedding] = await this.embeddingService.embedBatch([query]); const searchResult = await this.qdrantClient.search(this.collectionName, { vector: queryEmbedding, limit: topK, with_payload: true, }); return searchResult.map(result => ({ id: result.id.toString(), content: (result.payload as any).content, timestamp: (result.payload as any).timestamp, metadata: (result.payload as any).metadata, // score: result.score, // Could include score if needed })); } } ``` The collection is created with `size: 1536`, which has to match whatever the embedding model emits. Change the model string and forget the dimension and the store rejects everything you write to it. Inside the MoE agent, a `MemoryManager` expert decides *when* to store (after a significant observation, a successful tool execution, a complex reasoning step) and *what* to store, a summarised insight or the raw observation. It also fronts `search` for the other experts, so only one component knows the store exists. **The tell:** if you can't name the model string and the vector dimension your collection was created with, you have an abstraction, not a memory layer. --- ### Step 3 — Put the approval flag in the state An agent that needs human oversight has to pause, externalise the action it wants to take, wait, then proceed or revise. Every one of those verbs is a state transition, so approval belongs in the state object. ```typescript // types/agent.ts export interface AgentState { input: string; // User query scratchpad: string[]; // Internal monologue/reasoning steps tool_calls: { tool_name: string; args: Record; }[]; // Tools agent intends to call final_answer?: string; // If agent has a final answer awaitingHumanApproval: boolean; // Flag for HITL proposedAction?: { type: 'tool_call' | 'final_answer'; details: any; }; humanFeedback?: 'approve' | 'reject' | 'modify'; // What human decided modificationDetails?: string; // If human requested modification } // Initial state factory export const initialAgentState: () => AgentState = () => ({ input: '', scratchpad: [], tool_calls: [], awaitingHumanApproval: false, }); ``` `awaitingHumanApproval` is the pause. `proposedAction` is what the human is being asked about. `humanFeedback` has three values, not two, because *reject* and *modify* are different instructions to the agent's next think cycle. **The tell:** if the approval flag lives in a local variable inside a node, it cannot survive a checkpoint. The pause then dies with the process, and so does the pending decision. --- ### Step 4 — Make the pause a routing decision I lean toward raw APIs, but a branching agent workflow is one place a graph-based state machine earns its keep. `LangGraph` defines these well enough, verbose as it gets. Use it for state management and checkpointing, and keep the nodes themselves direct. The pieces: 1. **Agent State** — thoughts, observations, planned actions, and the approval flags. 2. **Workflow Graph** — the decision process, as a graph. 3. **Approval Node** — reached with a critical action, it flags `awaitingHumanApproval` and saves a checkpoint. 4. **External API** — a REST or WebSocket endpoint that receives the proposed action, shows it to a human, sends back the decision. 5. **Resume** — the state is updated with the decision and execution continues from the checkpoint. What makes the gate work is a single flag on the tool definition, `isCritical`, read at routing time: | tool | `isCritical` | route out of `agent_think` | |---|---|---| | `read_database` | `false` | `execute_tools` | | `deploy_code` | `true` | `await_human_approval` | The routing itself has a branch, a merge, and a cycle in it: ```mermaid flowchart TD T[agent think] --> G{route} G -->|critical action| P[await approval] G -->|tool calls| X[execute tools] G -->|final answer| E[END] P --> F[human feedback] F --> T X --> T ``` `LangGraph` integration needs `langchain_core` and `langgraph`. ```typescript // agentGraph.ts import { BaseMessage, HumanMessage, ToolMessage } from '@langchain/core/messages'; import { StateGraphArgs, END, StateGraph } from '@langchain/langgraph'; import { AgentState, initialAgentState } from './types/agent'; import { IVectorStoreClient } from './interfaces/memory'; import { QdrantVectorStoreClient } from './services/vectorStoreClient'; import { EmbeddingService } from './services/embeddingService'; import { ChatOpenAI } from '@langchain/openai'; // Using a direct LLM client // --- Mock Tools for demonstration --- interface Tool { name: string; description: string; schema: any; // Zod schema func: (args: any) => Promise; isCritical?: boolean; // New flag for HITL } const safeTool: Tool = { name: "read_database", description: "Reads non-sensitive data from a simulated database.", schema: { type: 'object', properties: { query: { type: 'string' } } }, func: async (args: { query: string }) => { console.log(`Executing safe tool: read_database with query: ${args.query}`); return `Data for "${args.query}" is: [itemA, itemB]`; }, isCritical: false, }; const criticalTool: Tool = { name: "deploy_code", description: "Deploys code to production. This is a critical operation.", schema: { type: 'object', properties: { project_id: { type: 'string' }, version: { type: 'string' } } }, func: async (args: { project_id: string; version: string }) => { console.warn(`CRITICAL: Deploying project ${args.project_id} version ${args.version}...`); return `Project ${args.project_id} version ${args.version} deployed successfully.`; }, isCritical: true, }; const tools: Tool[] = [safeTool, criticalTool]; // --- End Mock Tools --- const llm = new ChatOpenAI({ model: "gpt-4o-mini", // Or other performant model temperature: 0.7, }); // Bind tools to the LLM (LangChain specific) const llmWithTools = llm.bindTools(tools); // --- Agent Nodes --- // The main agent 'think' node (our MoE orchestrator/expert) const agentThinkNode = async (state: AgentState): Promise> => { console.log("[AGENT] Thinking..."); const messages: BaseMessage[] = [ new HumanMessage(state.input), ...state.scratchpad.map(s => new BaseMessage({ content: s, name: 'scratchpad_entry' })), // Example: add scratchpad to context ]; // Simulate retrieval from long-term memory for relevant context // In a real MoE, a 'MemoryExpert' would handle this. const relevantMemories = await memoryClient.search(state.input, 2); if (relevantMemories.length > 0) { messages.unshift(new HumanMessage(`Past relevant experiences: ${relevantMemories.map(m => m.content).join('\n')}`)); } const response = await llmWithTools.invoke(messages); const toolCalls = response.tool_calls || []; const content = response.content || ''; // Simulate the 'SafetyExpert' determining if an action is critical if (toolCalls.length > 0) { // Check if any proposed tool call is critical const criticalCall = toolCalls.find(tc => tools.find(t => t.name === tc.name)?.isCritical); if (criticalCall) { return { scratchpad: [...state.scratchpad, `Proposed critical tool call: ${JSON.stringify(criticalCall)}`], awaitingHumanApproval: true, proposedAction: { type: 'tool_call', details: criticalCall }, }; } } if (content.includes("FINAL ANSWER")) { // Simple heuristic for final answer return { scratchpad: [...state.scratchpad, content], final_answer: content.replace("FINAL ANSWER:", "").trim(), awaitingHumanApproval: true, // Also require approval for final answer if critical enough proposedAction: { type: 'final_answer', details: content }, }; } return { scratchpad: [...state.scratchpad, content], tool_calls: toolCalls.map(tc => ({ tool_name: tc.name, args: tc.args })), }; }; // Node to execute tools const executeToolsNode = async (state: AgentState): Promise> => { console.log("[AGENT] Executing tools..."); const toolResults: string[] = []; for (const toolCall of state.tool_calls) { const tool = tools.find(t => t.name === toolCall.tool_name); if (tool) { try { const result = await tool.func(toolCall.args); toolResults.push(`Tool ${toolCall.tool_name} returned: ${result}`); } catch (e) { toolResults.push(`Tool ${toolCall.tool_name} failed: ${e}`); } } else { toolResults.push(`Unknown tool: ${toolCall.tool_name}`); } } // Agent remembers the outcome of its actions await memoryClient.add([{ content: `Agent executed tools: ${state.tool_calls.map(tc => tc.tool_name).join(', ')}. Results: ${toolResults.join('; ')}`, timestamp: Date.now(), metadata: { type: 'tool_execution' } }]); return { scratchpad: [...state.scratchpad, ...toolResults], tool_calls: [], // Clear tool calls after execution }; }; // --- LangGraph Setup --- const graphState: StateGraphArgs['channels'] = { input: { value: (x: string, y: string) => y, // Overwrite default: () => '', }, scratchpad: { value: (x: string[], y: string[]) => x.concat(y), // Append default: () => [], }, tool_calls: { value: (x: any[], y: any[]) => y, // Overwrite default: () => [], }, final_answer: { value: (x: string | undefined, y: string | undefined) => y, // Overwrite default: () => undefined, }, awaitingHumanApproval: { value: (x: boolean, y: boolean) => y, // Overwrite default: () => false, }, proposedAction: { value: (x: any | undefined, y: any | undefined) => y, // Overwrite default: () => undefined, }, humanFeedback: { value: (x: 'approve' | 'reject' | 'modify' | undefined, y: 'approve' | 'reject' | 'modify' | undefined) => y, default: () => undefined, }, modificationDetails: { value: (x: string | undefined, y: string | undefined) => y, default: () => undefined, }, }; const workflow = new StateGraph({ channels: graphState, }) .addNode("agent_think", agentThinkNode) .addNode("execute_tools", executeToolsNode); // Define conditional edges for dynamic workflow workflow.addConditionalEdges( "agent_think", (state: AgentState) => { if (state.awaitingHumanApproval) { return "await_human_approval"; // Special state indicating pause } if (state.tool_calls.length > 0) { return "execute_tools"; } if (state.final_answer) { return END; } return "agent_think"; // Loop back to think if no tools/final answer } ); workflow.addConditionalEdges( "execute_tools", (state: AgentState) => { if (state.final_answer) { return END; } return "agent_think"; // After tools, go back to thinking } ); workflow.setEntryPoint("agent_think"); // Compile the graph const app = workflow.compile(); // Initialize memory client (global or passed into nodes) const embeddingSvc = new EmbeddingService(process.env.OPENAI_API_KEY!); const memoryClient = new QdrantVectorStoreClient(embeddingSvc, process.env.QDRANT_HOST || 'localhost:6333'); // --- External API for Human-in-the-Loop Management --- // This would typically be a separate microservice or API endpoint. // It retrieves the paused agent's state, presents it, and sends feedback. // Example of how human feedback would resume the agent export async function provideHumanFeedback(threadId: string, feedback: 'approve' | 'reject' | 'modify', modification?: string) { // In a real system, you'd load the current state of the thread/agent from LangGraph's checkpoint. // For demonstration, we'll simulate updating the state and resuming. // This is the critical part: LangGraph needs to be able to load the checkpoint // and continue execution with the updated state. // When LangGraph encounters a "return 'await_human_approval'" it implicitly // saves the state to its configured checkpointing mechanism (e.g., Redis, SQL). // To resume, you'd call `app.invoke` again with the thread_id and the // *updated* state for `humanFeedback` and `modificationDetails`. // LangGraph's checkpointing automatically handles loading the paused state. // We modify the *input* for the next step of the agent. const feedbackState: Partial = { awaitingHumanApproval: false, // Reset approval flag humanFeedback: feedback, modificationDetails: modification, // The agent's next 'think' cycle will process this feedback input: `Human feedback received: ${feedback}. ${modification ? `Modification requested: ${modification}` : ''}`, }; // This is where you would call LangGraph's `invoke` with the `threadId` // and the `feedbackState` as the updated state. LangGraph will then // load the checkpointed state for `threadId`, merge `feedbackState`, // and continue execution. // (Actual call depends on how LangGraph's checkpointing is configured and exposed) console.log(`[HITL] Human feedback for thread ${threadId}: ${feedback}. Modification: ${modification || 'none'}`); console.log("This feedback will now be fed back into the agent's workflow."); // In a production setup, you'd use a LangGraph `checkpoint_id` and resume // from there. For example: // const stream = app.stream(feedbackState, { // configurable: { thread_id: threadId } // }); // This would pick up from the last checkpoint and continue processing. } // --- Main execution flow (example) --- async function runAgentConversation(threadId: string, prompt: string) { let currentState: AgentState = initialAgentState(); currentState.input = prompt; console.log(`[USER] ${prompt}`); // LangGraph with configurable thread_id will manage checkpoints const stream = app.stream(currentState, { configurable: { thread_id: threadId }, }); for await (const s of stream) { if (s.__end__) { console.log("[AGENT] Workflow finished."); currentState = s.__end__ as AgentState; break; } const newNode = Object.keys(s).filter(key => key !== '__end__')[0]; console.log(`[GRAPH] Current node: ${newNode}`); currentState = s[newNode] as AgentState; if (currentState.awaitingHumanApproval) { console.warn(`[HITL REQUIRED] Agent proposed action: ${JSON.stringify(currentState.proposedAction)}. Waiting for human approval for thread ${threadId}...`); // Here, the system would pause. // A separate UI/API would pick up this state, show it to a human. // The human would then call `provideHumanFeedback(threadId, 'approve' | 'reject' | 'modify', modificationDetails)`. // The `stream` would then continue from the checkpoint. return currentState; // Exit here for simulation of waiting } } if (currentState.final_answer) { console.log(`[AGENT FINAL ANSWER] ${currentState.final_answer}`); } else { console.log("[AGENT] No final answer reached, current scratchpad:", currentState.scratchpad); } return currentState; } // Example Usage (requires setting OPENAI_API_KEY and QDRANT_HOST in .env or environment) // import { config } from 'dotenv'; // config(); // (async () => { // const thread1 = "user_123"; // let state; // // First interaction: non-critical tool // state = await runAgentConversation(thread1, "What is the capital of France? Also, read some database entries about 'users'."); // // Agent should run read_database and return answer // // Second interaction: critical tool // // state = await runAgentConversation(thread1, "Deploy project 'frontend-app' version '1.2.0' to production."); // // Agent should pause here, setting `awaitingHumanApproval: true` // // To simulate approval (after the above call returns with awaitingHumanApproval) // // if (state && state.awaitingHumanApproval) { // // console.log("Simulating human approval after 5 seconds..."); // // await new Promise(resolve => setTimeout(resolve, 5000)); // // await provideHumanFeedback(thread1, 'approve'); // // console.log("Resuming agent after human approval..."); // // await runAgentConversation(thread1, "Continue with deployment."); // Agent will pick up from checkpoint // // } // })(); ``` Two things in that file are doing more work than they look like they are. `agentThinkNode` pulls two memories from the vector store before it calls the model, so retrieval is part of thinking rather than a separate step. And `executeToolsNode` writes the outcome of every tool call back into memory, which is how the loop accumulates anything at all. **The tell:** if approval is an `if` statement inside a node instead of an edge out of it, the graph can't pause. It can only block, and a blocked worker is not a checkpoint. --- ### Step 5 — Resume from the checkpoint, not from the top `provideHumanFeedback` doesn't restart the agent. It builds a `Partial` that clears `awaitingHumanApproval`, records the decision, and rewrites `input` so the next think cycle reads the human's answer as its prompt. `LangGraph` loads the checkpoint for that `threadId`, merges the partial, and continues. The example above is simplified. In a real production system: - `LangGraph`'s checkpointing is configured with a persistent store, e.g. `RedisSupervisor`. - `app.stream` runs inside an asynchronous worker — a message queue consumer, not a request handler. - When `awaitingHumanApproval` goes true, the worker pushes the state to a queue or API that notifies the human approval system. - The approval system triggers `provideHumanFeedback`, which puts a message back on the worker's queue carrying the `thread_id` and the decision, and `app.stream` picks up from the last checkpoint. **The tell:** if resuming means replaying the user's original prompt, you aren't resuming. You're re-running, and every tool call before the gate happens twice. --- ### What breaks it **The framework tax.** `LangGraph` genuinely helps with a complex state machine, and the verbosity and overhead still grate. The rule that survives contact with the code: abstract where it helps, go down to the metal where performance or behaviour is critical. Embedding and vector search are on the metal side of that line, which is why Step 2 has no framework in it. **State across async steps with a human in the middle.** This is the hard part and it doesn't get easier. A mutable state object, an async worker, and a decision that might arrive in ten seconds or tomorrow. A clear `AgentState` and reliable checkpointing aren't optional here; they are the only thing holding the run together while nothing is executing. **Heuristics standing in for structure.** `content.includes("FINAL ANSWER")` decides whether the run terminates, and it's a string match on model output. `initializeCollection()` is fired without being awaited in the constructor, so the first `add` can race collection creation. Both are fine in a demo and both are the kind of thing that fails once, in production, on a day when nobody is reading logs. **The human factor.** Code is the smaller half of HITL. It needs a usable interface and a clear protocol for the operator, and the agent has to present its reasoning and its proposed action concisely enough that a human can decide quickly. Get that wrong and the approval gate becomes a queue nobody drains. --- ### When this architecture is the wrong choice The machinery costs latency, a vector store to operate, and a human whose time you are now spending. Four cases where it isn't worth it. - **The whole task fits in one session.** If nothing needs to survive the process, you don't need a hippocampus. You need a variable. - **You already know the key.** Semantic retrieval earns its cost when the query and the stored record share meaning but not words. Looking something up by an id it was filed under is a lookup, and an embedding call in front of it is latency you pay for nothing. - **The workflow doesn't branch.** A graph-based state machine earns its place on a branching workflow. On a straight line it's a verbose way to write a `for` loop, and the framework tax is the whole bill. - **No action in the system is expensive to undo.** The gate exists because deploys, emails, and payments are irreversible. If every tool is a read, `isCritical` is false everywhere and the approval path is dead code with a queue attached. --- ### The shift Building this stops being about chaining LLM calls and starts being about state management, data flow, and graceful failure. The interesting decisions all turned out to be about where state lives and who is allowed to change it. - **Give memory a boundary before you give it a backend.** `IVectorStoreClient` is two methods; the store behind it is replaceable. - **Make the pause a field on the state.** A flag that survives a checkpoint is a gate. A flag in a local variable is a blocked worker. - **Specialise the experts.** A `MemoryExpert` that owns storage and a `SafetyExpert` that owns the pause each do one job, and each can be tested alone. Comparing these systems to the brain keeps reminding me how early we are. Long-term memory and executive control are the baseline for anything you'd call intelligent, and what I want to try next is dynamic memory consolidation, learning from feedback loops, and adaptive MoE routing. --- Source: https://himanshuat.com/blogs/production-grade-agent-architecture-memory-human-in-the-loop ======================================================================== --- title: "Production-Grade RAG: A Blueprint for Scalable, Real-Time Architecture" description: "A RAG system fails quietly. It answers from a document that no longer exists, retrieval scores slide for weeks, and nothing in the trace is an error. This is the architecture I'd build to keep one fresh, fast, and observable." date: "November 27, 2025" url: "https://himanshuat.com/blogs/production-grade-rag-architecture-blueprint" --- # Production-Grade RAG: A Blueprint for Scalable, Real-Time Architecture The failure that kills a RAG system in production is not a crash. It is a confident answer assembled from a document that no longer exists, with no error anywhere in the trace. Getting a basic flow running with off-the-shelf libraries is easy. Holding up under load, staying fresh, keeping latency low, and not losing retrieval quality is a different problem from the prototype. > The prototype problem is retrieval quality. The production problem is freshness, and freshness is a property of the ingestion pipeline, not of the retriever. I'm drawn to how intelligence comes out of distributed, specialized systems, whether that's an MoE architecture or something more biological. Scaling RAG is partly an engineering problem and partly a way to think about that. The aim isn't to make it *work*, it's to make it *perform*, which means cutting anything in the path that doesn't earn its place. ### Chapter 0 — What the system actually is A production RAG system is four services, each owning one part of the flow: **ingestion**, the **vector database**, **retrieval**, and the **telemetry** that watches all three. The separation is the whole point. It's what lets you scale and maintain them independently, and it's why a bottleneck in one is diagnosable instead of ambient. Everything below assumes those four boundaries exist. If your RAG system is one process, most of this is premature. --- ### Step 1 — Push changes as events, not as a re-index A RAG system is only as good as the data it retrieves. In production, batch updates often aren't fresh enough. Stale information leads to hallucination or irrelevant answers, and periodic full re-indexing is slow. The fit is Change Data Capture or an event-driven approach pushed through a message queue. ```mermaid graph TD A["Source Systems (Databases, CMS, APIs)"] --> B("Change Data Capture / Webhooks") B --> C("Message Queue, e.g. Kafka") C --> D["Ingestion Service (Workers)"] D -- Embeddings --> E["Embedding Service (API)"] D -- Upsert / Delete --> F[Vector Database] D -- Metadata / Content --> G["Metadata Store (e.g. PostgreSQL, S3)"] ``` Three principles hold that together. Every content change — create, update, delete — emits an event. The queue decouples ingestion from the source systems, so the pipeline absorbs spikes instead of propagating them. And workers are idempotent, meaning the same message can be processed twice with no side effect, which is what makes retries safe. Skip the generic "document loaders" that frameworks ship. These are structured change events. Handling serialization directly is faster and keeps control over how data is transformed and embedded. ```typescript // src/ingestion-worker.ts import { Kafka } from 'kafkajs'; import { EmbeddingService } from './embeddingService'; // Raw API client for embeddings import { VectorDbClient } from './vectorDbClient'; // Raw API client for the vector DB interface DocumentChangeEvent { id: string; content: string; metadata: Record; type: 'upsert' | 'delete'; // Explicit event types: a delete is a first-class change timestamp: string; } const kafka = new Kafka({ brokers: [process.env.KAFKA_BROKERS || 'localhost:9092'] }); const consumer = kafka.consumer({ groupId: 'rag-ingestion-group' }); const embeddingService = new EmbeddingService(); const vectorDbClient = new VectorDbClient(); async function startIngestionWorker() { await consumer.connect(); await consumer.subscribe({ topic: 'document-changes', fromBeginning: false }); await consumer.run({ eachMessage: async ({ topic, partition, message, heartbeat, pause }) => { if (!message.value) { console.warn(`Received empty message on topic ${topic}, partition ${partition}`); return; } try { const event: DocumentChangeEvent = JSON.parse(message.value.toString()); console.log(`Processing event for docId: ${event.id}, type: ${event.type} at ${new Date().toISOString()}`); if (event.type === 'upsert') { const embedding = await embeddingService.getEmbedding(event.content); // Upsert, not insert: replaying the same event twice must not duplicate a point. await vectorDbClient.upsert({ id: event.id, vector: embedding, metadata: { ...event.metadata, ingestion_timestamp: event.timestamp }, content: event.content // Store raw content for LLM context, or reference a separate store }); console.log(`Successfully upserted document ${event.id}`); } else if (event.type === 'delete') { await vectorDbClient.delete(event.id); console.log(`Successfully deleted document ${event.id}`); } await heartbeat(); // Keep consumer session alive } catch (error) { console.error(`Error processing message from topic ${topic}, partition ${partition}, offset ${message.offset}:`, error); // Dead-letter queue, metrics, alerts belong here. A swallowed error is a silently stale index. // pause(); // Example: Pause consumption if error rate is too high. } }, }); } startIngestionWorker().catch(err => { console.error("Ingestion worker failed to start:", err); process.exit(1); }); ``` The delete branch is the one people leave out, and it's the one that produces answers citing documents that were removed. **The tell:** if the only way a new document reaches your index is the next scheduled indexing job, your freshness ceiling is that interval, and no amount of retrieval tuning moves it. --- ### Step 2 — Choose a vector database for filtering and sharding, not for the demo The vector database is the core store, holding the semantic representation of your knowledge base. Skip in-memory toy databases in production. **Qdrant**, **Milvus**, **Weaviate**, **Pinecone**, and **Vespa** give you horizontal scaling for high-throughput reads and writes, replication for fault tolerance, filtering for precise retrieval by user role or document type or date, and low-latency search via HNSW indexing. With the right DevOps, self-hosted options like Qdrant can beat managed services on cost at scale. Scaling it is four levers, and they buy different things: | lever | what it buys | |---|---| | Shard across nodes | capacity | | Replicate across nodes | read-heavy throughput and availability | | Tune `M`, `ef_construction`, `ef_search` | latency matched to your dataset size and target | | Quantize | a smaller memory footprint, for a small recall trade-off | Wrapping the database in an ORM-like layer adds overhead in the hottest path you have. Direct API calls keep it fast. ```typescript // src/vectorDbClient.ts import { QdrantClient } from '@qdrant/qdrant-sdk'; import { PointStruct, Filter, RecommendRequest } from '@qdrant/qdrant-sdk/dist/qdrant_client'; export class VectorDbClient { private client: QdrantClient; private collectionName: string; constructor( host: string = process.env.QDRANT_HOST || 'localhost', port: number = parseInt(process.env.QDRANT_PORT || '6333', 10), collectionName: string = process.env.QDRANT_COLLECTION || 'rag_documents' ) { this.client = new QdrantClient({ host, port }); this.collectionName = collectionName; this.ensureCollectionExists().catch(console.error); } private async ensureCollectionExists() { const { collections } = await this.client.getCollections(); const collectionExists = collections.some(c => c.name === this.collectionName); if (!collectionExists) { console.log(`Creating Qdrant collection: ${this.collectionName}`); await this.client.createCollection(this.collectionName, { vectors: { size: 1536, distance: 'Cosine' }, // Assuming OpenAI text-embedding-3-small (1536 dim) // on_disk_payload, replication_factor, etc. belong here for production }); console.log(`Collection ${this.collectionName} created.`); } } async upsert(document: { id: string; vector: number[]; metadata: Record; content?: string }) { const point: PointStruct = { id: document.id, vector: document.vector, payload: { ...document.metadata, content: document.content || null } // Store content in payload }; await this.client.upsert(this.collectionName, { wait: true, // Wait for operation to be finished points: [point], }); } async query( queryVector: number[], limit: number = 5, filters?: Filter // Qdrant's explicit Filter type: precision happens here, not in the prompt ): Promise }>> { const searchResult = await this.client.search(this.collectionName, { vector: queryVector, limit: limit, filter: filters, with_payload: true, // Always retrieve payload (metadata + content) with_vectors: false, // Vectors are dead weight on the response path }); return searchResult.map(hit => ({ id: hit.id, score: hit.score, payload: hit.payload || {}, })); } async delete(id: string) { await this.client.delete(this.collectionName, { pointsSelector: { points: [id] } }); } // `recommend` exists for advanced retrieval patterns built on positive/negative examples async recommend( positiveVectors: number[][], negativeVectors: number[][] = [], limit: number = 5, filters?: Filter ): Promise }>> { const request: RecommendRequest = { positive: positiveVectors, negative: negativeVectors, limit: limit, filter: filters, with_payload: true, }; const recommendResult = await this.client.recommend(this.collectionName, request); return recommendResult.map(hit => ({ id: hit.id, score: hit.score, payload: hit.payload || {}, })); } } ``` **The tell:** if you have never changed `ef_search`, you are shipping whatever latency-recall trade-off the default happened to pick for you. --- ### Step 3 — Give every cache layer its own TTL This is where a query becomes an answer. The retrieval service embeds the query, searches the vector database with metadata filters for precision (`source_type: 'internal_docs'`), builds a tight prompt around what came back, and calls the LLM. ```mermaid graph TD A[User Query] --> B(API Gateway) B --> C[Retrieval Service] C -- 1. Generate Query Embedding --> D["Embedding Service (API)"] C -- 2. Query Vector DB (with filters) --> E[Vector Database] C -- 3. Construct Prompt --> F["LLM Service (API)"] F --> C C -- 4. Return Response --> B C -- Caching --> G[Redis Cache] ``` Caching isn't an afterthought here. It's part of the design, for latency and for cost, and the three layers do not age at the same rate: | layer | keyed on | TTL | |---|---|---| | query embedding | the query text | 24 hours | | retrieved documents | query, filters, result count | 300s | | final LLM response | query, retrieved doc ids, filters | 120s | An embedding for a fixed query is stable, so it can sit for a day. Retrieval results go stale at the speed of your ingestion pipeline. A final response is the most expensive thing to regenerate and the fastest thing to invalidate, because any new document changes what the right answer is. LangChain's chains and agents tend to hide the actual flow. I want explicit control over each step — embedding, retrieval, prompt construction, the LLM call — and caching that's granular rather than a wrapper around everything. ```typescript // src/retrievalService.ts import { EmbeddingService } from './embeddingService'; import { VectorDbClient } from './vectorDbClient'; import { LLMClient } from './llmClient'; import { RedisClient } from './redisClient'; import crypto from 'crypto'; // For generating cache keys export class RetrievalService { private embeddingService: EmbeddingService; private vectorDbClient: VectorDbClient; private llmClient: LLMClient; private redis: RedisClient; constructor() { this.embeddingService = new EmbeddingService(); this.vectorDbClient = new VectorDbClient(); this.llmClient = new LLMClient(); this.redis = new RedisClient(); } // Stable hash so the same query always lands on the same key private generateCacheKey(input: string): string { return crypto.createHash('sha256').update(input).digest('hex'); } // Generic caching helper: one TTL per call site, never one global TTL private async getCached(key: string, fetchFn: () => Promise, ttlSeconds: number = 3600): Promise { const cached = await this.redis.get(key); if (cached) { // console.log(`Cache hit for ${key}`); // Uncomment for debugging cache hits return JSON.parse(cached) as T; } // console.log(`Cache miss for ${key}`); // Uncomment for debugging cache misses const result = await fetchFn(); await this.redis.set(key, JSON.stringify(result), ttlSeconds); return result; } async retrieveAndGenerate( userQuery: string, filters?: Record, // Dynamic filters from API numDocuments: number = 5 ): Promise { const startTime = Date.now(); // 1. Generate Query Embedding (and cache it) const embeddingCacheKey = this.generateCacheKey(`embedding:${userQuery}`); const queryEmbedding = await this.getCached( embeddingCacheKey, () => this.embeddingService.getEmbedding(userQuery), 3600 * 24 // Embeddings for the same query are stable, cache longer ); // logMetric('embedding_latency_ms', Date.now() - startTime, 'histogram', { type: 'query' }); // 2. Retrieve Documents from Vector DB (and cache the retrieval results) const filterString = filters ? JSON.stringify(filters) : ''; const retrievalCacheKey = this.generateCacheKey(`retrieval:${userQuery}:${filterString}:${numDocuments}`); const retrievedDocs = await this.getCached( retrievalCacheKey, async () => { const hits = await this.vectorDbClient.query(queryEmbedding, numDocuments, filters); return hits.map(hit => ({ content: hit.payload.content as string, metadata: hit.payload as Record })); }, 300 // Retrieval results might become stale faster, depending on ingestion speed ); // logMetric('vector_db_retrieval_latency_ms', Date.now() - startTime, 'histogram'); if (retrievedDocs.length === 0) { console.warn("No relevant documents found for query. Using LLM without context."); // Fallback: answer without context. Count these; they are ungrounded answers. return await this.llmClient.generateResponse([{ role: 'user', content: userQuery }]); } // 3. Construct LLM Prompt const context = retrievedDocs.map(doc => doc.content).join('\n\n---\n\n'); const promptMessages = [ { role: 'system', content: 'You are a helpful assistant. Use the provided context to answer the user\'s question. If the answer is not explicitly in the context, state that you don\'t know, but do not make up information.' }, { role: 'user', content: `Context:\n${context}\n\nQuestion: ${userQuery}` } ]; // 4. Generate LLM Response (and cache the final response) // Key includes the doc ids, so a different retrieval set is a different answer const finalResponseCacheKey = this.generateCacheKey( `response:${userQuery}:${retrievedDocs.map(d => d.metadata.id).sort().join(',')}:${filterString}` ); const finalResponse = await this.getCached( finalResponseCacheKey, () => this.llmClient.generateResponse(promptMessages), 120 // Final responses are very expensive, but can become stale quickly with new data ); // logMetric('llm_generation_latency_ms', Date.now() - startTime, 'histogram'); // logMetric('rag_total_latency_ms', Date.now() - startTime, 'histogram', { cache_hit: finalResponseWasCached ? 'true' : 'false' }); return finalResponse; } } // Minimal supporting clients for context: // src/llmClient.ts import OpenAI from 'openai'; import { ChatCompletionMessageParam } from 'openai/resources/chat/completions'; export class LLMClient { private openai: OpenAI; constructor() { this.openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY }); } async generateResponse(messages: ChatCompletionMessageParam[]): Promise { try { const completion = await this.openai.chat.completions.create({ model: process.env.LLM_MODEL || 'gpt-4o-mini', // Model is config, not code messages: messages, temperature: 0.2, // Keep temperature low for factual RAG }); return completion.choices[0].message.content || ''; } catch (error) { console.error("Error calling LLM:", error); throw new Error("Failed to generate response from LLM."); } } } // src/embeddingService.ts export class EmbeddingService { private openai: OpenAI; // Re-use OpenAI client for embeddings if using OpenAI models constructor() { this.openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY }); } async getEmbedding(text: string): Promise { try { const response = await this.openai.embeddings.create({ model: process.env.EMBEDDING_MODEL || 'text-embedding-3-small', input: text, }); return response.data[0].embedding; } catch (error) { console.error("Error generating embedding:", error); throw new Error("Failed to generate embedding."); } } } // src/redisClient.ts import { createClient, RedisClientType } from 'redis'; export class RedisClient { private client: RedisClientType; private isConnected: boolean = false; constructor(url: string = process.env.REDIS_URL || 'redis://localhost:6379') { this.client = createClient({ url }); this.client.on('error', (err) => { console.error('Redis Client Error', err); this.isConnected = false; }); this.client.on('connect', () => { console.log('Connected to Redis'); this.isConnected = true; }); this.client.connect().catch(err => console.error("Failed to connect to Redis on startup:", err)); } async get(key: string): Promise { if (!this.isConnected) return null; // A dead cache degrades latency, never correctness try { return await this.client.get(key); } catch (error) { console.error(`Error getting from Redis key ${key}:`, error); return null; } } async set(key: string, value: string, ttlSeconds: number): Promise { if (!this.isConnected) return; try { await this.client.set(key, value, { EX: ttlSeconds }); } catch (error) { console.error(`Error setting to Redis key ${key}:`, error); } } } ``` Notice what the Redis client does on failure: it returns `null` and the request goes to the real service. A cache that can take down the request path is not a cache, it's a dependency. **The tell:** if all three of your cache layers share one TTL, at least two of them are wrong. --- ### Step 4 — Instrument for silent degradation, not just for errors If you can't measure it, you can't improve it. Own your telemetry rather than leaning on a black-box SaaS for the metrics that actually matter, because production RAG has plenty of subtle failure modes that quietly degrade answers without ever raising one. | layer | what you watch | |---|---| | Ingestion | queue depth, producer/consumer lag, worker latency, failed embeddings and writes | | Vector database | query latency, throughput, index size, memory, recall against ground truth | | Cache | hit rate per layer, cache size, eviction rate | | LLM service | API latency, success rate, input/output token usage for cost | | End to end | total request latency, retrieval service error rate | Prometheus and Grafana for metrics and alerts, a centralized stack (ELK, Grafana Loki, Datadog) for structured logs, OpenTelemetry for tracing a request across all four services. ```typescript // src/telemetry.ts type MetricType = 'counter' | 'gauge' | 'histogram'; export function logMetric(name: string, value: number, type: MetricType, labels?: Record) { // In production this exposes metrics on an HTTP endpoint Prometheus scrapes // (e.g. prom-client in Node.js) rather than writing to stdout. const labelStr = labels ? Object.entries(labels).map(([k, v]) => `${k}=${v}`).join(',') : ''; console.log(`METRIC: ${name}{${labelStr}} ${value} (type: ${type}, timestamp: ${Date.now()})`); // Example for a Prometheus client: // if (type === 'histogram') { // promClient.histogram(name, labels).observe(value); // } else if (type === 'counter') { // promClient.counter(name, labels).inc(value); // } // ... } // Example usage in retrievalService.ts: // ... inside retrieveAndGenerate method ... // const startTime = performance.now(); // Use performance.now() for high-res timing // ... call embeddingService ... // logMetric('rag_embedding_latency_ms', performance.now() - startTime, 'histogram', { source: 'user_query' }); // ... after vector DB query ... // logMetric('rag_vector_db_query_latency_ms', performance.now() - vectorDbStartTime, 'histogram', { hits: retrievedDocs.length > 0 ? 'true' : 'false' }); // ... after cache check ... // logMetric('rag_cache_hit_rate', cacheHitCount / (cacheHitCount + cacheMissCount), 'gauge', { cache_layer: 'retrieval_results' }); ``` **The tell:** if your dashboards show only errors and latency, you will hear about retrieval degradation from a user rather than from a graph. --- ### Four drift signals to watch this week RAG quality degrades silently as the data distribution shifts, the embedding model ages, or the index fragments. These four are independent, and none of them show up as an error: - **Retrieval score distribution.** Track similarity scores for retrieved documents. A sudden drop in the average or median points to stale data, a weak embedding model, or an index problem. - **Source and metadata distribution.** Check where retrieved documents come from. Are you always pulling from old sources? Are important document types underrepresented? - **Null retrieval rate.** How often does a query return nothing relevant? A high rate means gaps in the knowledge base or poor query understanding. - **Human feedback.** Sample responses periodically for human review, then correlate those scores against the automated metrics. That correlation is what tells you whether the automated metrics are worth watching at all. --- ### What breaks it **The cache outlives the data.** Nothing in the ingestion path evicts anything from Redis. A delete event removes the point from the vector database, and a cached response built from that point keeps serving until its TTL expires on its own. The TTLs are the entire invalidation strategy, which is why the response layer gets 120 seconds and not an hour. **The empty-retrieval fallback hides the failure.** When retrieval returns nothing, the service calls the LLM with no context and returns the answer through the same interface as a grounded one. It looks identical to the caller. That's what makes the null retrieval rate a metric rather than a log line — it's the only place the ungrounded answers are visible. **Drift has no error to raise.** A score distribution that slides by a few percent a week never trips a threshold anybody set. You catch it by having recorded what normal looked like before it moved, which means the baseline has to exist before the degradation starts. --- ### When this architecture is the wrong choice Every piece of this costs operational surface. Four services, a broker, a cache tier, and a metrics stack are all things that can page you at night. - **Your corpus barely changes.** The entire case for CDC and a message queue is freshness. If the underlying documents update monthly, a scheduled re-index is slower and simpler, and simpler wins. - **You have no latency or cost pressure yet.** Three cache layers with three separate TTLs and three invalidation stories exist to buy back milliseconds and API spend. If nobody is complaining about either, you're maintaining a cache for the pleasure of it. - **You have no ground truth.** Recall metrics need evaluations to measure against, and drift detection needs a recorded baseline. Without either, the observability layer produces graphs nobody can act on. Build the evaluation set first. - **You're still discovering the shape of the flow.** The argument for raw API clients over framework abstractions is control, and control is only valuable once you know what you want controlled. While the pipeline is still changing weekly, the framework's defaults are cheaper than your hand-rolled ones. --- ### The shift The things that made this work are all unglamorous. Events instead of batches. A separate TTL per cache layer. A metric on the path where the system fails quietly. None of it is about the retriever, which is where almost all the prototype tuning goes. - **Make freshness a pipeline property.** Every change emits an event, deletes included, and workers are idempotent so retries are free. - **Give each cache layer its own TTL and its own staleness story.** One global TTL means at least two layers are wrong. - **Instrument the silent path.** Score distribution, null retrieval rate, and source mix, because none of them will ever throw. The next problems I want to work on are dynamic re-ranking, post-retrieval processing that summarizes chunks before they reach the LLM, self-healing ingestion, and active learning that fine-tunes the embedding and re-ranking models from real usage. > A system that retrieves information is a starting point. The interesting version understands it and presents it well, and none of that is reachable until the boring layer underneath stops lying to you. --- Source: https://himanshuat.com/blogs/production-grade-rag-architecture-blueprint ======================================================================== --- title: "Part 5: Production-Ready Agents - Implementing Human-in-the-Loop Supervision" description: "An agent that can execute a $100,000 transfer needs a state that means wait. This is a small TypeScript orchestrator where the pause is a persisted state rather than an exception, and resuming means writing the human's decision into state before the loop runs again." date: "April 26, 2024" url: "https://himanshuat.com/blogs/production-ready-langgraph-human-in-the-loop" --- # Part 5: Production-Ready Agents - Implementing Human-in-the-Loop Supervision An agent reads a request, decides it means transfer $100,000 to account X, and executes it. Nothing in the run stopped, because nothing in the run could stop. The model was working from a picture of the world, the picture was wrong, and the workflow had no state that meant *wait for a person*. > Full autonomy in a critical domain is not a capability milestone. It is a missing state in your state machine. A modular, MoE-inspired architecture buys you resilience and specialised expertise. It does not buy you a pause. Financial transactions, medical diagnoses, infrastructure management: in those domains a single hallucination is expensive, and the fix is not a better expert. It's a place in the graph where execution stops and a human decides. Human-in-the-loop is what closes the gap between an impressive prototype and something that survives contact with an ambiguous input. It validates critical decisions before they execute, corrects outputs when they drift, produces explicit signal for future model improvements, and satisfies regulations that require human oversight. ### Chapter 0 — What an interruption actually is An interruption is not a blocked thread, a callback, or a promise nobody resolves. Those all die with the process. **An interruption is a state the workflow can be persisted in, and that no node is allowed to advance past without an external trigger.** Three pieces carry it. `WorkflowState` is a structured object written to a durable store — PostgreSQL JSONB, Redis, DynamoDB. It holds: * `workflowId: string` * `currentState: AgentState` (`INITIALIZING`, `PROCESSING_STEP_1`, `WAITING_FOR_HUMAN`, `COMPLETED`, `FAILED`) * `data: Record` — the mutable context passed between nodes * `lastExecutedNode: string | null` * `interruptionDetails?: { nodeId: string; reason: string; humanInputRequired: boolean }` `AgentNode` is one discrete piece of agent logic: an async function from state to state, and nothing else. ```typescript type AgentNodeFn = (state: WorkflowState) => Promise; ``` `AgentOrchestrator` is the engine. It loads the current state, picks the next node from a definition, executes it, persists the result, and stops when the state says stop. That is the whole vocabulary. No framework required. --- ### Step 1 — Persist the state, not the call stack Most agent frameworks abstract state management and orchestration away. They get you started fast, and they charge for it in overhead, imposed structure, and lost granular control. For an interruptible workflow that tradeoff is bad, because the thing you most need to control is exactly the thing that got abstracted. So the requirements are explicit: state with no hidden globals and no magically passed context, low serialisation and function-call overhead, interruption points anywhere rather than only at predefined tool calls, and a logged, traceable record of every transition. A workflow that pauses for a human is a workflow that may sit paused for hours. If the pause lives in a stack frame, a deploy kills it. **The tell:** if killing the process that holds a paused workflow loses the pause, you have a call stack, not a state machine. --- ### Step 2 — Let the node raise its own pause The condition that triggers human review is business logic. It belongs in the node that knows the business, not in the engine. `analyzeRequestNode` below runs its analysis, checks the result, and if the work looks critical it flips `currentState` to `await_human_review` and fills in `interruptionDetails`. The orchestrator never learns what "critical" means. ```typescript // types.ts interface WorkflowState { workflowId: string; currentState: string; // e.g., 'start', 'analyze_request', 'propose_action', 'await_human_review', 'execute_action', 'complete' data: Record; // Dynamic data for the workflow lastExecutedNode: string | null; interruptionDetails?: { nodeId: string; reason: string; humanInputRequired: boolean; }; } // nodes.ts const initialRequestNode: AgentNodeFn = async (state) => { console.log(`[${state.workflowId}] Initializing request...`); // Example: parse initial input, set up context state.data.initialQuery = state.data.inputQuery; // Assume inputQuery comes from initial trigger state.currentState = 'analyze_request'; state.lastExecutedNode = 'initialRequestNode'; return state; }; const analyzeRequestNode: AgentNodeFn = async (state) => { console.log(`[${state.workflowId}] Analyzing request: ${state.data.initialQuery}`); // Use an LLM or specific expert to analyze // const analysisResult = await callLLMExpert(state.data.initialQuery); state.data.analysis = "LLM analysis results here..."; // Placeholder // Simple conditional for interruption AFTER this node if (state.data.analysis.includes("critical financial operation")) { console.log(`[${state.workflowId}] Critical operation detected. Requiring human review.`); state.currentState = 'await_human_review'; state.interruptionDetails = { nodeId: 'analyzeRequestNode', reason: 'Critical operation identified, requires human approval', humanInputRequired: true }; } else { state.currentState = 'propose_action'; } state.lastExecutedNode = 'analyzeRequestNode'; return state; }; const proposeActionNode: AgentNodeFn = async (state) => { // This node could be interrupted BEFORE execution if human already intervened // (e.g., human explicitly marked 'no_action_needed' in previous step) if (state.interruptionDetails && state.interruptionDetails.humanInputRequired && !state.data.humanApproved) { // This is a "guard" for interrupt_before functionality // The orchestrator would typically handle this, but for clarity, showing it here. throw new Error("Workflow is paused, human input pending or disapproved."); } console.log(`[${state.workflowId}] Proposing action based on analysis: ${state.data.analysis}`); // const actionPlan = await generateActionPlan(state.data.analysis); state.data.actionPlan = "Generated action plan here..."; // Placeholder state.currentState = 'execute_action'; state.lastExecutedNode = 'proposeActionNode'; return state; }; const executeActionNode: AgentNodeFn = async (state) => { if (!state.data.humanApproved) { // This could be another 'interrupt_before' scenario if the human explicitly rejected // and the orchestrator didn't update the state correctly. throw new Error("Action not approved by human. Cannot execute."); } console.log(`[${state.workflowId}] Executing action: ${state.data.actionPlan}`); // await runAction(state.data.actionPlan); // External API call state.data.executionResult = "Action completed successfully."; // Placeholder state.currentState = 'complete'; state.lastExecutedNode = 'executeActionNode'; return state; }; const completionNode: AgentNodeFn = async (state) => { console.log(`[${state.workflowId}] Workflow completed.`); state.currentState = 'complete'; state.lastExecutedNode = 'completionNode'; return state; }; // Map node names to functions const nodes: Record = { initialRequestNode, analyzeRequestNode, proposeActionNode, executeActionNode, completionNode }; ``` Note what `proposeActionNode` and `executeActionNode` do at the top: they refuse to run unless `state.data.humanApproved` is set. Those guards are the other half of the mechanism, and they matter more than they look. Covered in Step 4. **The tell:** if adding a new approval gate means editing the orchestrator, the gate condition is in the wrong file. --- ### Step 3 — Loop on state, exit on interruption The orchestrator does five things: load the current state, pick the next node from `currentState`, execute it, persist the new state, and, if `currentState` signals a pause, stop instead of advancing. The loop condition is the entire safety property. It runs while the workflow is not complete, not failed, and has no `interruptionDetails`. ```typescript // orchestrator.ts import { WorkflowState, AgentNodeFn } from './types'; // Assuming types defined above interface WorkflowDefinition { [stateName: string]: { node: string; // Name of the node function nextState: string | ((state: WorkflowState) => string); // Next state or a function to determine it }; } // A simple in-memory store for demonstration. In production, use Redis/DB. const workflowStore = new Map(); class AgentOrchestrator { private workflowDefinition: WorkflowDefinition; private nodeFunctions: Record; constructor(workflowDefinition: WorkflowDefinition, nodeFunctions: Record) { this.workflowDefinition = workflowDefinition; this.nodeFunctions = nodeFunctions; } private async saveState(state: WorkflowState): Promise { workflowStore.set(state.workflowId, { ...state }); // Deep copy to prevent mutation issues // console.log(`[${state.workflowId}] State saved: ${state.currentState}`); } private async loadState(workflowId: string): Promise { return workflowStore.get(workflowId); } public async initializeWorkflow(workflowId: string, initialData: Record): Promise { const initialState: WorkflowState = { workflowId, currentState: 'start', // Entry point data: initialData, lastExecutedNode: null, }; await this.saveState(initialState); return initialState; } public async execute(workflowId: string): Promise { let state = await this.loadState(workflowId); if (!state) { throw new Error(`Workflow ${workflowId} not found.`); } while (state.currentState !== 'complete' && state.currentState !== 'failed' && !state.interruptionDetails) { const currentStep = this.workflowDefinition[state.currentState]; if (!currentStep) { throw new Error(`Invalid state in workflow definition: ${state.currentState}`); } const nodeFn = this.nodeFunctions[currentStep.node]; if (!nodeFn) { throw new Error(`Node function '${currentStep.node}' not found.`); } try { // Execute the node state = await nodeFn(state); // If the node itself set an interruption, break the loop if (state.interruptionDetails) { console.log(`[${state.workflowId}] Workflow paused for human review at node: ${state.lastExecutedNode}`); await this.saveState(state); return state; } // Determine next state if (typeof currentStep.nextState === 'function') { state.currentState = currentStep.nextState(state); } else if (currentStep.nextState) { state.currentState = currentStep.nextState; } else { // Implicitly move to next state if not defined, or assume completion // For a robust system, this should be explicitly handled in definition. state.currentState = 'complete'; } await this.saveState(state); } catch (error: any) { console.error(`[${workflowId}] Error executing node ${currentStep.node}:`, error.message); state.currentState = 'failed'; state.data.error = error.message; await this.saveState(state); return state; } } console.log(`[${workflowId}] Workflow execution finished. Current state: ${state.currentState}`); return state; } // This is the external API endpoint handler for resuming a workflow public async resumeWithHumanInput(workflowId: string, humanInput: Record): Promise { let state = await this.loadState(workflowId); if (!state) { throw new Error(`Workflow ${workflowId} not found.`); } if (!state.interruptionDetails) { throw new Error(`Workflow ${workflowId} is not currently interrupted.`); } // Apply human input to the workflow data state.data = { ...state.data, ...humanInput }; // Clear interruption status state.interruptionDetails = undefined; // Decide the next step based on human input or resume from where it left off // For simplicity, we'll assume human input dictates approval and we resume from the next logical step // In a real system, humanInput might include a 'decision' field like 'approved'/'rejected' if (humanInput.decision === 'approved') { state.data.humanApproved = true; // Flag for subsequent nodes // Resume from the node that would have been next if no interruption const interruptedNodeConfig = this.workflowDefinition[state.currentState]; if (typeof interruptedNodeConfig.nextState === 'function') { state.currentState = interruptedNodeConfig.nextState(state); } else { state.currentState = interruptedNodeConfig.nextState || 'complete'; // Or whatever makes sense } console.log(`[${workflowId}] Human approved. Resuming from state: ${state.currentState}`); } else { // Human rejected, potentially move to a 'rejected' or 'failed' state state.data.humanApproved = false; state.currentState = 'failed'; state.data.error = 'Human rejected the proposed action.'; console.log(`[${workflowId}] Human rejected. Workflow moved to state: ${state.currentState}`); } await this.saveState(state); return this.execute(workflowId); // Attempt to continue execution } } // Define the workflow's state transitions const agentWorkflowDefinition: WorkflowDefinition = { 'start': { node: 'initialRequestNode', nextState: 'analyze_request' }, 'analyze_request': { node: 'analyzeRequestNode', nextState: (state) => state.interruptionDetails ? 'await_human_review' : 'propose_action' }, 'await_human_review': { // This is a "waiting" state, orchestrator won't execute further node: 'analyzeRequestNode', // This node already ran, this just signifies we're paused AFTER it nextState: (state) => state.data.humanApproved ? 'propose_action' : 'failed' // After human input }, 'propose_action': { node: 'proposeActionNode', nextState: 'execute_action' }, 'execute_action': { node: 'executeActionNode', nextState: 'complete' }, 'complete': { node: 'completionNode', nextState: 'complete' } // Terminal state }; // --- Example Usage --- async function runExample() { const orchestrator = new AgentOrchestrator(agentWorkflowDefinition, nodes); const workflowId = `wf-${Date.now()}`; console.log("--- Starting Workflow 1 (requires human approval) ---"); let state1 = await orchestrator.initializeWorkflow(workflowId, { inputQuery: "Perform critical financial operation: Transfer $100,000 to account X" }); state1 = await orchestrator.execute(workflowId); console.log("\nCurrent state after initial run:", state1.currentState, state1.interruptionDetails); if (state1.currentState === 'await_human_review' && state1.interruptionDetails) { console.log("\n--- Simulating human review and approval ---"); // An external API call would hit `resumeWithHumanInput` // For example: POST /api/workflow/{workflowId}/resume { decision: 'approved', comment: 'Looks good' } state1 = await orchestrator.resumeWithHumanInput(workflowId, { decision: 'approved', comment: 'Review passed.' }); console.log("State after human approval and resumption attempt:", state1.currentState); } console.log("\n--- Starting Workflow 2 (no human approval needed) ---"); const workflowId2 = `wf-${Date.now()}-no-interrupt`; let state2 = await orchestrator.initializeWorkflow(workflowId2, { inputQuery: "Generate a summary of recent tech news" }); state2 = await orchestrator.execute(workflowId2); console.log("State after full execution (no interrupt):", state2.currentState); } runExample(); ``` `agentWorkflowDefinition` is the whole graph, and it fits on a screen. Two conditional edges, one waiting state, one terminal state, one failure state. ```mermaid flowchart TD A[start] --> B[analyze request] B -->|routine| C[propose action] B -->|critical| D{await human} D -->|approved| C D -->|rejected| F[failed] C --> E[execute action] E --> G[complete] ``` **The tell:** read your loop condition. If it doesn't mention the interruption field, the loop will run straight through a paused workflow. --- ### Step 4 — Resume by writing the decision in before you re-enter the loop `resumeWithHumanInput` is the external API handler. A reviewer POSTs a decision, and the handler merges the human input into `state.data`, clears `interruptionDetails`, sets `state.data.humanApproved`, computes the next state from the workflow definition, saves, and only then calls `execute` again. The ordering is the point. State is updated *before* the loop runs, so the guards at the top of `proposeActionNode` and `executeActionNode` see the decision rather than racing it. That gives you two distinct gate placements out of one mechanism. | gate | where the check runs | what stops the run | |---|---|---| | `interrupt_after` | inside `analyzeRequestNode`, after its work completes | the node sets `interruptionDetails`; the orchestrator sees it, persists, and returns | | `interrupt_before` | at the top of `proposeActionNode` and `executeActionNode` | the guard throws unless `state.data.humanApproved` is set | There is no standalone `interrupt_before` flag here, and there could be. Adding one to `WorkflowDefinition` and having the orchestrator check it *before* calling `nodeFn` would move that guard out of the node and into the engine, where the check is uniform and impossible to forget in a new node. **The tell:** if your resume endpoint calls the next node directly instead of updating state and handing back to the loop, you now have two execution paths, and they will drift. --- ### Three things to build this week The code above runs. It does not yet hold traffic. 1. **A real store.** `workflowStore` is a `Map`. Swap it for a key-value store like Redis for fast state lookups, or an indexed table if you want `WorkflowState` in a relational database. 2. **Idempotent nodes.** A node that gets retried after a failure should not repeat its side effects. This matters most for `executeActionNode`, the one node that touches the outside world. 3. **Transition logging.** State transitions, node execution times, and payload sizes. Auditability was one of the reasons for building this instead of importing it, and you only get it if you write the log lines. --- ### What breaks it **Two writers, one workflow.** Nothing here prevents a race. If a resume call and a background executor both load, mutate, and save the same `WorkflowState`, the last write wins and the other decision disappears. Production needs transactional updates or optimistic locking, and it needs them specifically around `saveState`. **A guard that throws is a workflow that dies.** The `interrupt_before` guards in `proposeActionNode` and `executeActionNode` throw. The orchestrator's `catch` turns any thrown error into `currentState = 'failed'`. So a run that reaches those nodes with approval still pending doesn't pause politely, it terminates with `Workflow is paused, human input pending or disapproved.` in `state.data.error`. Pausing and failing should not share an exit path. **The waiting state points at a node that already ran.** `await_human_review` maps to `analyzeRequestNode` because the definition requires every state to name a node. The only thing keeping that node from running a second time is the `interruptionDetails` check in the loop condition. Clear the flag wrong and the analysis re-runs, re-detects the critical operation, and re-pauses forever. --- ### When human-in-the-loop is the wrong choice The pause is not free. It converts one run into two runs with human latency in between, and it costs somebody's attention every time it fires. - **The action is cheap and reversible.** The whole argument rests on catastrophic consequences: money moved, a diagnosis issued, infrastructure changed. Generating a summary of recent tech news is the second example in the code for a reason, and it never pauses. - **Nobody is on the other end.** A paused workflow waits for an external `resumeWithHumanInput` call. Without a reviewer, a queue, and an endpoint they actually use, an interruption is not supervision. It's a stall with better logging. - **You cannot express the gate condition.** The trigger here is a check on the analysis text. If you cannot write the predicate that separates "needs a human" from "doesn't", you cannot place the gate, and gating everything is the same as having no agent. - **You don't need the granular control.** Building the orchestrator is the right call when you need explicit state, interruption points anywhere, and low overhead. If none of those are load-bearing for you, a framework's opinionated structure is cheaper than maintaining an engine yourself. --- ### The shift Rolling this by hand instead of importing a framework was worth the extra code, for the same reason raw API calls beat an ORM on a critical path: the messy parts are exactly the parts you need to tune. The larger reframe is that we got graph-like state transitions without a graph library. A state machine, nodes and edges, expresses cleanly in plain code and runs leaner. - **Make the pause a state, not an exception.** States persist across restarts. Stack frames don't. - **Put the trigger in the node and the stop in the loop.** The node knows what critical means; the engine only needs to know that something is set. - **Update state before you resume, never after.** The guards downstream read state, so state has to be true by the time they run. Human oversight is a feature, not an admission that the AI failed. Build the gate before you need it, because the run that needs it will not ask first. --- Source: https://himanshuat.com/blogs/production-ready-langgraph-human-in-the-loop ======================================================================== --- title: "The Promise and Potential of Quantum Computing" description: "An overview of quantum computing, its principles, applications, and future prospects." date: "February 05, 2023" url: "https://himanshuat.com/blogs/quantum-computing-simplified" --- # The Promise and Potential of Quantum Computing Quantum computing uses quantum-mechanical phenomena, such as superposition and entanglement, to perform operations on data. These phenomena let a quantum computer tackle some problems far more efficiently than a classical one. ![Quantum Computing](https://unsplash.com/photos/WdJkXFQ4VHY) _Photo by Manuel on Unsplash_ A traditional computer uses bits, units of information that can only be 0 or 1. A quantum computer uses qubits, which can be 0, 1, or both at once (superposition). Qubits can also be entangled, meaning the state of one qubit depends on the state of another. This allows a quantum computer to perform many calculations at the same time, which is known as parallel computing. A traditional computer can only perform one calculation at a time. ## Applications of Quantum Computing ### Quantum Cryptography A well-known application is quantum cryptography, which uses the laws of quantum mechanics to encrypt information. The encryption is considered unbreakable and is used to secure sensitive data such as banking transactions and government communications. ### Quantum Simulation Quantum simulation is the ability to simulate quantum systems on a quantum computer. It can be used to study systems like molecules and materials and understand their properties. ### Artificial Intelligence Quantum computing may also help with artificial intelligence, since it can be used to train and run machine learning algorithms. ## The Future of Quantum Computing Despite its many advantages, quantum computing is still in its early stages of development, and it will take many more years before it becomes widely available. Quantum computing could change how we approach some hard computations, secure information, and model quantum systems, with potential impact in fields like healthcare, finance, and transportation. It's still early, though, and a long way from being widely available. --- Source: https://himanshuat.com/blogs/quantum-computing-simplified ======================================================================== --- title: "Part 3: Scaling LangGraph - State Persistence, Checkpointing, and Parallelism" description: "A long-running graph crashes and loses the reasoning chain that was mid-flight, because nothing in it was ever written down. Checkpointers and Pregel supersteps fix two different halves of that, and neither one gives you the parallelism you probably think you are buying." date: "May 03, 2024" url: "https://himanshuat.com/blogs/scaling-langgraph-persistence-and-parallelism" --- # Part 3: Scaling LangGraph - State Persistence, Checkpointing, and Parallelism A long-running agent crashes partway through a run. The reasoning chain that was in flight is gone, and so is every turn of context it accumulated getting there. Nothing in the graph noticed, because there was nothing in the graph to notice with. A stateless agent is a conversation with someone who forgets everything you said five minutes ago. > Persistence and parallelism aren't features you bolt onto a graph. They're properties of its structure, and you commit to both the moment you draw the edges. My long-term obsession is replicating parts of a human brain with a Modular-by-Design (MoE) architecture. The brain isn't one monolithic process. It's a parallel network of specialized modules that fire, talk to each other, and remember, and every thought is state transitions plus information moving between them. My earlier LangGraph prototypes (Part 1, Part 2) were fun proofs of concept, and they were linear, short-lived, and stateless. Even a simplified brain simulation has to survive crashes, run for a long time, and think about more than one thing at once. That isn't a smarter-LLM problem. It's a system that has to hold context indefinitely and process information efficiently. So there are two things to get: **state persistence**, which is memory, and **parallelism**, which is concurrent processing. LangGraph has an abstraction for each. --- ### Chapter 0 — What a checkpointer actually is A **checkpointer** snapshots the entire state of your graph to an external store. Input, current node, pending messages, outputs. To resume, the graph loads that snapshot and picks up exactly where it left off. That is the whole primitive. Everything else is a question of where the snapshot lands and how much it costs to put it there. The API is convenient, and my competitive-programmer instinct still asks the uncomfortable version of the question. *How* is the state stored? *How* is it retrieved? What's the overhead? For a serious MoE system I'd build a custom state layer eventually, probably on something like Agno for low-latency, type-safe access in a distributed setup. Understanding the shipped abstractions is how you learn where you'd drop below them. --- ### Step 1 — Start on `SqliteSaver`, and know why you'll leave it For local development or a single-instance app, `SqliteSaver` is the cheapest thing that works. It serializes your graph's state into a SQLite database file. The thing to watch in this snippet is `thread_id`. It is the handle the checkpointer keys on, and it is what makes the second `invoke` a continuation rather than a fresh start. ```python import os from langgraph.graph import StateGraph, START from langgraph.checkpoint.sqlite import SqliteSaver # Define our simple graph state (accumulator for demonstration) class GraphState: messages: list turn: int # Define a simple node that appends a message and increments a turn counter def chat_node(state: GraphState): current_turn = state.turn new_message = f"Agent turn {current_turn}: Processing..." return {"messages": state.messages + [new_message], "turn": current_turn + 1} # Initialize our checkpointer memory = SqliteSaver.from_conn_string(":memory:") # Use an in-memory SQLite DB for ephemeral testing # Or use a file for persistence: # memory = SqliteSaver.from_conn_string("sqlite:///checkpoints.sqlite") # Build the graph workflow = StateGraph(GraphState) workflow.add_node("chat", chat_node) workflow.set_entry_point("chat") workflow.add_edge("chat", "chat") # Self-loop for continuous chat app = workflow.compile(checkpointer=memory) # Define initial state thread_id = "test_thread_1" initial_state = {"messages": ["User: Hello!"], "turn": 0} # Run for a few steps print(f"--- Running thread '{thread_id}' for the first time ---") for i in range(3): print(f"Step {i+1}") output = app.invoke( input=initial_state if i == 0 else {}, # Only pass initial state on first invoke config={"configurable": {"thread_id": thread_id}} ) print(output['messages'][-1]) # The output from invoke already includes the latest state, but we can also inspect it directly # from the checkpointer if we were to restart the app. # Simulate application restart or new process print("\n--- Simulating restart and resuming ---") # If using a file-based sqlite, just re-initialize the app with the same checkpointer path. # For in-memory, we can just invoke again as it's still in the same process memory. # Load the state for the same thread_id # If we were truly restarting, we'd initialize `app` again. # The next invoke will automatically load the last state for `test_thread_1` output_resumed = app.invoke( input={}, # No new input, just continue from previous state config={"configurable": {"thread_id": thread_id}} ) print(f"Resumed output after restart: {output_resumed['messages'][-1]}") print(f"Total messages: {len(output_resumed['messages'])}") # Run another step print("\n--- Running another step after resumption ---") output_final = app.invoke( input={}, config={"configurable": {"thread_id": thread_id}} ) print(f"Final output: {output_final['messages'][-1]}") print(f"Total messages: {len(output_final['messages'])}") ``` The second `invoke` passes an empty input and still produces the next message in the sequence. That's the checkpointer doing its job. `SqliteSaver` is a non-starter for distributed systems. It's single-file, not built for concurrent writes, and it doesn't give you the low-latency, high-availability behavior a production MoE agent needs. Anything approaching a brain needs fast, shared memory. **The tell:** if the same `thread_id` in a new process doesn't resume mid-conversation, you have a database with no memory in it. The store is wired but the key isn't. --- ### Step 2 — Move the snapshot to Redis once more than one process needs it `RedisSaver` is a real step up. Redis is an in-memory data structure store, good for caching and persisting state with low latency, which is closer to what a distributed MoE system needs for shared state. The graph is unchanged. Only the store swaps. ```python import os import redis from langgraph.graph import StateGraph, START from langgraph.checkpoint.redis import RedisSaver # Assuming Redis is running on localhost:6379 # (Or set REDIS_URL environment variable) # e.g., os.environ["REDIS_URL"] = "redis://localhost:6379/0" # Verify Redis connection try: _ = redis.Redis.from_url(os.getenv("REDIS_URL", "redis://localhost:6379/0")).ping() print("Successfully connected to Redis.") except redis.exceptions.ConnectionError as e: print(f"Error connecting to Redis: {e}") print("Please ensure Redis is running or set REDIS_URL environment variable.") exit(1) # Reuse the same GraphState and chat_node from above class GraphState: messages: list turn: int def chat_node(state: GraphState): current_turn = state.turn new_message = f"Agent turn {current_turn}: Processing..." return {"messages": state.messages + [new_message], "turn": current_turn + 1} # Initialize our Redis checkpointer # It will use REDIS_URL env var or default to redis://localhost:6379/0 memory = RedisSaver() # Build and compile the graph (same as before) workflow = StateGraph(GraphState) workflow.add_node("chat", chat_node) workflow.set_entry_point("chat") workflow.add_edge("chat", "chat") app_redis = workflow.compile(checkpointer=memory) thread_id_redis = "test_thread_redis_1" initial_state_redis = {"messages": ["User: Hello via Redis!"], "turn": 0} print(f"\n--- Running Redis-backed thread '{thread_id_redis}' ---") for i in range(3): print(f"Step {i+1}") output = app_redis.invoke( input=initial_state_redis if i == 0 else {}, config={"configurable": {"thread_id": thread_id_redis}} ) print(output['messages'][-1]) print("\n--- Simulating restart and resuming Redis-backed thread ---") # In a real scenario, this would be a new process. # Just re-invoking will fetch state from Redis. output_resumed_redis = app_redis.invoke( input={}, config={"configurable": {"thread_id": thread_id_redis}} ) print(f"Resumed output after restart (Redis): {output_resumed_redis['messages'][-1]}") print(f"Total messages (Redis): {len(output_resumed_redis['messages'])}") # Cleanup (optional): Clear state for this thread_id # If you want to manually inspect the Redis keys, you can do: # redis_client = redis.Redis.from_url(os.getenv("REDIS_URL", "redis://localhost:6379/0")) # keys_to_delete = redis_client.keys(f"langgraph:{thread_id_redis}:*") # if keys_to_delete: # redis_client.delete(*keys_to_delete) # print(f"Cleaned up {len(keys_to_delete)} keys for thread {thread_id_redis} in Redis.") ``` The two savers differ on exactly the axes that decide whether you can put a second process in front of the graph. | checkpointer | store | concurrent writers | fits | |---|---|---|---| | `SqliteSaver` | single local file | not built for it | local iteration, single instance | | `RedisSaver` | in-memory, networked | yes | distributed shared state | Redis is the better production choice for its speed and distributed support, and `RedisSaver` is still an abstraction. For peak performance and fine-grained control I'd talk to Redis keys directly and write my own serialization, especially for custom data types that don't serialize well out of the box. The competitive programmer in me wants to know the exact network round trips and byte sizes. **The tell:** if two processes need to read the same thread and exactly one of them owns the file on disk, you have already outgrown `SqliteSaver`. The migration isn't a performance decision at that point, it's a correctness one. --- ### Step 3 — Draw the graph so independent work has no edge between it A brain doesn't process information sequentially. Billions of neurons fire in parallel, and the patterns they form add up to coherent thought. LangGraph's engine is built on the **Pregel** model, a graph processing framework designed for iterative, vertex-centric computation. Computation proceeds in a series of *supersteps*. In each superstep a vertex receives messages from the previous superstep, computes, updates its state, and sends messages onward. The consequence that matters: **nodes without direct dependencies run in parallel within a superstep.** There is no `Thread.start()` and no `asyncio.gather()` in LangGraph's top-level `invoke`. The structure of your edges decides what can proceed independently, which means concurrency is something you draw, not something you call. Here is the shape the next snippet builds. One planner fans out to three researchers, and all three merge back into one writer. ```mermaid flowchart LR P[plan] --> A[research ethics] P --> B[research quantum] P --> C[research neuroscience] A --> W[write report] B --> W C --> W W --> E[END] ``` The three research nodes have no edge between them. That absence is the entire parallelism story. ```python from typing import List, Literal, TypedDict from langgraph.graph import StateGraph, END import time # Define a more complex state for demonstration class MultiAgentState(TypedDict): research_topics: List[str] report_sections: List[str] final_report: str current_step: Literal["plan", "research", "write", "finish"] # Node 1: Plan topics (sequential) def plan_topics(state: MultiAgentState): print(f"[{time.time():.2f}] Planner: Starting to plan topics...") time.sleep(0.5) # Simulate work topics = ["AI Ethics", "Quantum Computing", "Neuroscience in AI"] print(f"[{time.time():.2f}] Planner: Planned topics: {topics}") return {"research_topics": topics, "current_step": "research"} # Node 2a: Research AI Ethics (can run in parallel with 2b, 2c) def research_ai_ethics(state: MultiAgentState): if "AI Ethics" in state["research_topics"]: print(f"[{time.time():.2f}] Researcher A: Starting 'AI Ethics' research...") time.sleep(1.0) # Simulate heavy research section = "AI Ethics: Discusses fairness, bias, and responsible AI development." print(f"[{time.time():.2f}] Researcher A: Finished 'AI Ethics'.") return {"report_sections": [section]} return {"report_sections": []} # Node 2b: Research Quantum Computing (can run in parallel) def research_quantum_computing(state: MultiAgentState): if "Quantum Computing" in state["research_topics"]: print(f"[{time.time():.2f}] Researcher B: Starting 'Quantum Computing' research...") time.sleep(0.8) # Simulate heavy research section = "Quantum Computing: Explores superposition, entanglement, and QML applications." print(f"[{time.time():.2f}] Researcher B: Finished 'Quantum Computing'.") return {"report_sections": [section]} return {"report_sections": []} # Node 2c: Research Neuroscience in AI (can run in parallel) def research_neuroscience_ai(state: MultiAgentState): if "Neuroscience in AI" in state["research_topics"]: print(f"[{time.time():.2f}] Researcher C: Starting 'Neuroscience in AI' research...") time.sleep(1.2) # Simulate heavy research section = "Neuroscience in AI: Examines brain-inspired algorithms and neural networks." print(f"[{time.time():.2f}] Researcher C: Finished 'Neuroscience in AI'.") return {"report_sections": [section]} return {"report_sections": []} # Node 3: Write Final Report (sequential, depends on all research) def write_report(state: MultiAgentState): print(f"[{time.time():.2f}] Editor: Starting to compile report...") time.sleep(0.7) # Simulate writing combined_sections = "\n\n".join(state["report_sections"]) final_report_content = f"Comprehensive Report:\n\n{combined_sections}\n\nConclusion: ... Generated on {time.strftime('%Y-%m-%d')}" print(f"[{time.time():.2f}] Editor: Finished compiling report.") return {"final_report": final_report_content, "current_step": "finish"} # Build the graph for parallel execution workflow_parallel = StateGraph(MultiAgentState) # Add nodes workflow_parallel.add_node("plan", plan_topics) workflow_parallel.add_node("research_ethics", research_ai_ethics) workflow_parallel.add_node("research_quantum", research_quantum_computing) workflow_parallel.add_node("research_neuroscience", research_neuroscience_ai) workflow_parallel.add_node("write_report", write_report) # Define entry point and edges workflow_parallel.set_entry_point("plan") # After planning, all research nodes can potentially run workflow_parallel.add_edge("plan", "research_ethics") workflow_parallel.add_edge("plan", "research_quantum") workflow_parallel.add_edge("plan", "research_neuroscience") # All research nodes must complete before writing the report. # LangGraph's Pregel engine will manage waiting for all incoming edges to complete. workflow_parallel.add_edge("research_ethics", "write_report") workflow_parallel.add_edge("research_quantum", "write_report") workflow_parallel.add_edge("research_neuroscience", "write_report") # End the graph workflow_parallel.add_edge("write_report", END) app_parallel = workflow_parallel.compile() # Initial state initial_state_parallel = MultiAgentState( research_topics=[], report_sections=[], final_report="", current_step="plan" ) print("\n--- Running Parallel Research Graph ---") final_output_parallel = app_parallel.invoke(initial_state_parallel) print("\n--- Final Report ---\n") print(final_output_parallel['final_report']) ``` Every node prints a timestamp on entry and exit, which is the only instrumentation you need to check the claim. In the output, `research_ethics`, `research_quantum`, and `research_neuroscience` start very close to each other, *after* `plan` completes. The engine found the independent paths from the edges alone, and it held `write_report` until all three incoming edges landed. **The tell:** trace every pair of nodes in your graph. If every pair is joined by a directed path, the graph is a line wearing a diagram's clothes, and no superstep will ever hold more than one node. --- ### Step 4 — Put the real parallelism inside the node Pregel orchestrates the steps. It does not make the work inside a step concurrent. For CPU-bound or I/O-bound parallelism inside a single node, say fetching several APIs at once, you still need Python's `asyncio`, `multiprocessing`, or `concurrent.futures`. LangGraph structures the high-level flow. Node-level optimization is yours. The distinction worth holding onto is between two kinds of parallelism that look identical on a diagram. | kind | who provides it | what it operates on | |---|---|---| | logical | LangGraph's Pregel engine | independent nodes in a superstep | | physical | `asyncio`, `multiprocessing`, `concurrent.futures` | work inside one node | For the MoE vision, where hundreds or thousands of expert modules might run at once, I'd eventually swap LangGraph's execution layer for a distributed task queue, Celery with Redis or a custom gRPC worker pool, to get real inter-process parallelism. **The tell:** delete LangGraph from the picture and look at one node. If the work inside it would still need `asyncio.gather` to go faster, that speedup was never LangGraph's to give you. --- ### What breaks it **Concurrent writes to a single file.** `SqliteSaver` is one file and isn't built for concurrent writers. The moment a second process holds the same graph, the checkpoint store becomes the contention point, and it fails in the least visible place possible: state. **Serialization you didn't choose.** `RedisSaver` decides how your state becomes bytes. Custom data types that don't serialize cleanly out of the box are where the abstraction stops being free, and the fix is to talk to Redis keys directly and own the serialization yourself. **Confusing the diagram with the execution.** A fan-out in the graph means LangGraph will schedule those nodes in the same superstep. It says nothing about what happens inside them. A node that blocks is a node that blocks, and no amount of edge-drawing changes that. --- ### When LangGraph's abstractions are the wrong choice Every layer of abstraction costs something. These are the cases where the cost isn't repaid. - **The run is short, linear, and disposable.** A checkpointer buys crash recovery and resumption. If the whole thing finishes in one process and you'd rerun it rather than resume it, the store is machinery around a problem you don't have. - **Every node depends on the one before it.** Pregel finds parallelism in the absence of edges. In a chain there is no absence to find, and the superstep model is bookkeeping you're paying for and never using. - **The bottleneck is inside a node, not between them.** CPU-bound work across cores, or concurrent I/O, lives in the node implementation. Reaching for the graph to fix it is reaching for the wrong layer. - **You need real inter-process parallelism at scale.** LangGraph is a good pattern for graph state machines, not a distributed computing framework. Hundreds or thousands of concurrent expert modules is a task queue's job, Celery with Redis or a gRPC worker pool, not a graph executor's. - **You need control over serialization, indexing, and transactional guarantees.** That's the point at which you write a custom low-latency key-value layer, or an Agno-backed state layer, and stop asking the checkpointer to guess. --- ### The shift Persistence isn't optional for anything long-running, and the choice of store is a system design decision rather than a config line. `RedisSaver` is a practical step toward a distributed system. For a production MoE simulation I'd reach for a custom low-latency key-value store or an Agno-backed state layer, so I control serialization, indexing, and transactional guarantees. Understanding Pregel changes how you draw graphs, not just how you read them. Once you know supersteps schedule whatever has no edge between it, you start looking for independent paths on purpose and cutting the bottlenecks that force everything back into a line. - **Pick the store before you need it.** Migrating a checkpointer under a live thread is a correctness problem, not a tuning one. - **Draw the independence.** Parallelism you didn't put in the edges doesn't exist. - **Fix the node with node tools.** `asyncio` for I/O, `multiprocessing` for CPU, and stop asking the graph. Ideally I'd write a lean graph executor from scratch, maybe in TypeScript with Agno for type safety and low-level control, so every cycle goes into the simulation. Until then the honest position is that LangGraph gives you useful patterns, and brain-like fidelity will push you below them. --- Source: https://himanshuat.com/blogs/scaling-langgraph-persistence-and-parallelism ======================================================================== --- title: "The Self-Correcting RAG: Implementing Agentic and Recursive Retrieval Loops" description: "Static RAG retrieves once and hopes the first search was enough. This is the loop version, where the model judges its own context, names what is missing, writes a new query, and searches again until it stops or hits the cap. TypeScript, raw APIs, ideas from FLARE and Self-RAG." date: "November 18, 2025" url: "https://himanshuat.com/blogs/self-correcting-rag-agentic-recursive-retrieval" --- # The Self-Correcting RAG: Implementing Agentic and Recursive Retrieval Loops The first search returns five documents. The model writes a fluent answer. The answer is wrong in the one place those five documents were silent, and nothing in the pipeline was looking for that hole. That's the standard shape: embed a query, retrieve, cram the results into a prompt, synthesise. A single shot that assumes the first retrieval was both complete and relevant. > Retrieval isn't the hard part. Noticing that the retrieval was insufficient is the hard part, and a single-shot pipeline has nowhere to put that judgment. That's not how you work through a hard problem either. You don't recall one thing and stop. You notice what you're missing, ask a follow-up, search again, and cross-reference until it fits together. So the goal stops being *retrieve* and becomes *reason about retrieval*: find the gaps, write new queries, build the answer up recursively. Ideas from FLARE and Self-RAG, no heavy framework hiding the mechanics. TypeScript, raw APIs, direct control. ### Chapter 0 — What a retrieval loop actually is A retrieval loop is a recursive agent with four stages. Start with the user query, embed it, and run a first search. The agent then reads the original query against everything retrieved so far and decides whether the context answers the query fully, what's missing if it doesn't, and what new search query would fill that gap. If more is needed, run the new query and add the results to the accumulated context. Once the agent signals it's done, the model writes the final answer from everything gathered. ```mermaid flowchart TD Q[user query] --> E[embed + search] E --> C[accumulate context] C --> J{context enough?} J -->|NEEDS_MORE_INFO| F[write follow-up query] F --> E J -->|COMPLETE| S[synthesise] ``` The edge back from the decision node to the search node is the whole post. Everything else is plumbing that exists to make that edge safe to traverse. --- ### Step 1 — Make retrieval something the loop can call twice Retrieval has to be a function, not a stage. Whether the store is Pinecone, Weaviate, Qdrant, or a local FAISS index, it reduces to embedding a query and fetching nearest neighbours, so the interface stays small enough to call from inside a `while`. ```typescript // For brevity, we'll abstract away the specific Vector DB client // In a real project, this would wrap your specific client (e.g., Pinecone, Weaviate) interface VectorDB { search: (embedding: number[], k: number) => Promise; } interface Document { id: string; content: string; metadata?: Record; embedding?: number[]; // Potentially store embeddings, or generate on-the-fly } // A simple abstraction for an embedding service const getEmbedding = async (text: string): Promise => { // Example using OpenAI's embedding API directly const response = await fetch('https://api.openai.com/v1/embeddings', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': `Bearer ${process.env.OPENAI_API_KEY}`, }, body: JSON.stringify({ input: text, model: 'text-embedding-ada-002', // Or a more performant model like 'text-embedding-3-small' }), }); if (!response.ok) { throw new Error(`Embedding API error: ${response.statusText}`); } const data = await response.json(); return data.data[0].embedding; }; ``` Note that `search` takes `k` as an argument rather than baking it in. The loop passes a different `k` on the first pass than on follow-ups, because a follow-up is targeting a named gap and doesn't need the same width. **The tell:** if your retrieval function is called exactly once per request, you don't have a loop. You have a query. --- ### Step 2 — Make the model's decision parseable, not readable The loop's control flow depends on what the model says, which means the model's output has to be a structure and not a paragraph. A direct call with `response_format: { type: "json_object" }` is the whole mechanism. ```typescript type LLMCallOptions = { model: string; temperature?: number; max_tokens?: number; }; // Generic LLM call function const callLLM = async (messages: Array<{ role: 'system' | 'user' | 'assistant', content: string }>, options: LLMCallOptions): Promise => { // Example for OpenAI API const response = await fetch('https://api.openai.com/v1/chat/completions', { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': `Bearer ${process.env.OPENAI_API_KEY}`, }, body: JSON.stringify({ model: options.model, messages: messages, temperature: options.temperature || 0.7, max_tokens: options.max_tokens, response_format: { type: "json_object" } // Crucial for structured output }), }); if (!response.ok) { throw new Error(`LLM API error: ${response.statusText}`); } const data = await response.json(); return data.choices[0].message.content; }; ``` One function serves both the decision call and the final synthesis call. The difference between them is the prompt and the `max_tokens`, which is what lets you swap models per role later without touching the transport. **The tell:** if you're pattern-matching the model's prose for the word "complete", your control flow is a string match on a paragraph. --- ### Step 3 — Ask what's missing, not whether the model is satisfied A yes/no from the model gives the loop nothing to search with. The decision prompt asks for three fields: a `status`, a `critique` naming the gap, and a `followUpQuery` targeting it. Previous searches go in too, so the agent doesn't spend an iteration re-running a query it already ran. ```typescript interface AgentDecision { status: 'COMPLETE' | 'NEEDS_MORE_INFO'; followUpQuery?: string; critique?: string; } const getAgentDecisionPrompt = (originalQuery: string, currentContext: string[], accumulatedSearchHistory: string[], maxTokens: number): Array<{ role: 'system' | 'user', content: string }> => [ { role: 'system', content: `You are an expert research agent. Your task is to evaluate a user's query against provided context. Your goal is to determine if the current context is sufficient to answer the original query comprehensively. If it is, output status: 'COMPLETE'. If not, you must identify the knowledge gaps, provide a 'critique' of what's missing, and formulate a precise 'followUpQuery' to find the missing information. The follow-up query should be specific and target the identified gap. You have a maximum context window of ${maxTokens} tokens for your final answer. Manage it wisely. Output your decision as a JSON object with 'status', 'followUpQuery' (if status is 'NEEDS_MORE_INFO'), and 'critique' (if status is 'NEEDS_MORE_INFO').` }, { role: 'user', content: `Original Query: "${originalQuery}" --- Currently Accumulated Context: ${currentContext.length > 0 ? currentContext.map((doc, i) => `Document ${i + 1}:\n${doc}`).join('\n---\n') : 'No context yet.'} --- Previous Search History (to avoid redundant searches): ${accumulatedSearchHistory.length > 0 ? accumulatedSearchHistory.join('\n') : 'No previous searches.'} --- Based on the above, is the current context sufficient? If not, what should the next search query be? JSON Output:` } ]; const evaluateContext = async ( originalQuery: string, currentContext: string[], accumulatedSearchHistory: string[], llmOptions: LLMCallOptions, maxTokensForFinalAnswer: number ): Promise => { const prompt = getAgentDecisionPrompt(originalQuery, currentContext, accumulatedSearchHistory, maxTokensForFinalAnswer); const responseJson = await callLLM(prompt, llmOptions); try { return JSON.parse(responseJson) as AgentDecision; } catch (error) { console.error("Failed to parse agent decision JSON:", responseJson, error); // Fallback: If LLM screws up JSON, assume it needs more info with a generic query return { status: 'NEEDS_MORE_INFO', followUpQuery: `Refine search for "${originalQuery}" based on current context.`, critique: "LLM failed to output valid JSON. Assuming more info needed." }; } }; ``` The `critique` field is the one that earns its keep. It's not decoration for the prompt: it's the only record of *why* the loop went around again, and when a follow-up query comes back useless, the critique is where you find out whether the model misread the gap or read it correctly and searched badly. **The tell:** if your decision object has a status and no critique, you can debug that the loop iterated but never why. --- ### Step 4 — Bound the loop before you run it Two things go wrong in a loop like this, and the constants exist to stop both. Token windows blow out as context accumulates, and the loop runs forever when the model never says `COMPLETE`. | knob | value | what it bounds | |---|---|---| | `MAX_RAG_ITERATIONS` | 3 | total recursive searches before the loop exits regardless | | `CONTEXT_CHUNK_SIZE` | 4000 | tokens per context chunk before summarisation or truncation | | `FINAL_ANSWER_MAX_TOKENS` | 1500 | the synthesis budget, also told to the decision agent | | `k_initial` | 5 | documents pulled by the first, broad search | | `k_followup` | 3 | documents pulled by each targeted follow-up | Those numbers are defaults, not measurements. What matters is that every one of them is a ceiling the model doesn't get a vote on. ```typescript const MAX_RAG_ITERATIONS = 3; // Limit the number of recursive searches const CONTEXT_CHUNK_SIZE = 4000; // Max tokens per context chunk before summarization/truncation const FINAL_ANSWER_MAX_TOKENS = 1500; // Max tokens for the final generated answer interface RecursiveRAGResult { answer: string; iterations: number; finalContext: string[]; searchHistory: string[]; } const runRecursiveRAG = async ( userQuery: string, vectorDB: VectorDB, llmOptions: LLMCallOptions, k_initial: number = 5, // Number of docs for initial search k_followup: number = 3 // Number of docs for follow-up searches ): Promise => { let accumulatedContext: string[] = []; let searchHistory: string[] = []; let currentIteration = 0; let currentSearchQuery = userQuery; while (currentIteration < MAX_RAG_ITERATIONS) { currentIteration++; console.log(`--- Iteration ${currentIteration}: Searching for "${currentSearchQuery}" ---`); searchHistory.push(currentSearchQuery); const embedding = await getEmbedding(currentSearchQuery); const docs = await vectorDB.search(embedding, currentIteration === 1 ? k_initial : k_followup); const newContexts = docs.map(d => d.content); // Simple context management: Add new context, deduplicate, and maybe truncate if too large // For production, this would involve more sophisticated summarization or re-ranking accumulatedContext = Array.from(new Set([...accumulatedContext, ...newContexts])); // Prompt LLM to evaluate context and decide next steps const decision = await evaluateContext( userQuery, accumulatedContext, searchHistory, llmOptions, FINAL_ANSWER_MAX_TOKENS ); if (decision.status === 'COMPLETE') { console.log(`Agent decided context is COMPLETE. Critque: ${decision.critique || 'N/A'}`); break; // Exit loop, context is deemed sufficient } else { console.log(`Agent needs more info. Critique: ${decision.critique}. Follow-up Query: "${decision.followUpQuery}"`); currentSearchQuery = decision.followUpQuery!; // LLM must provide a query if status is 'NEEDS_MORE_INFO' } } // Final synthesis step console.log("--- Final Synthesis ---"); const finalSynthesisPrompt = [ { role: 'system', content: `You are an expert answer generator. Based on the provided context, answer the original query comprehensively and concisely. Ensure your answer directly addresses all aspects of the query. If the context is insufficient for a part of the query, state that explicitly. Your final answer should not exceed ${FINAL_ANSWER_MAX_TOKENS} tokens.` }, { role: 'user', content: `Original Query: "${userQuery}" --- Context for Synthesis: ${accumulatedContext.map((doc, i) => `Document ${i + 1}:\n${doc}`).join('\n---\n')} --- Please provide the most comprehensive answer possible based on the above context.` } ]; const finalAnswer = await callLLM(finalSynthesisPrompt, { ...llmOptions, max_tokens: FINAL_ANSWER_MAX_TOKENS }); return { answer: finalAnswer, iterations: currentIteration, finalContext: accumulatedContext, searchHistory: searchHistory }; }; ``` Two details in there are load-bearing. The synthesis prompt is instructed to say so explicitly when the context is insufficient for part of the query, which means a loop that ran out of iterations degrades into an honest partial answer rather than a confident invented one. And the result carries `iterations`, `finalContext`, and `searchHistory` back out, so a bad answer is inspectable after the fact instead of being a single opaque string. **The tell:** if the only thing that ends your loop is the model saying `COMPLETE`, the model's worst day is your bill. --- ### Step 5 — Split the model across the loop The decision call and the synthesis call are not the same job. The decision call answers a narrow structured question and runs several times per query. The synthesis call runs once and produces the thing the user reads. So use a smaller, faster model for the decision step, something in the `gpt-3.5-turbo` class or an equivalent from Agno or elsewhere, and save the more capable model for the final synthesis. Across three iterations that split is cheaper and quicker for the same behaviour. That split is a micro-MoE of sorts: routing by role rather than by topic, with the expensive expert reserved for the step that needs it. **The tell:** if the model writing your final answer is also the one answering a yes/no three times per query, you're paying synthesis prices for a classification. --- ### Three things to build this week These are independent of each other. Pick whichever the current failure points at. - **Context compaction before the evaluation call.** As `accumulatedContext` grows you hit token limits. Summarise documents or groups of them before passing them on, re-rank with a smaller model or a heuristic against the *original* query and *current* gaps and drop the weak ones, or window it by passing the most recent `N` documents plus a summary of the rest. - **Parallel follow-ups for independent gaps.** If the agent spots several gaps that don't depend on each other, the follow-up queries could run at the same time. That needs an agent that can decompose the problem, which is a bigger change than it sounds. - **A visible thought process.** The search history and the per-iteration `critique` are already a rough record of how the system reasoned. Surfacing that internal monologue helps behaviour and debugging both. --- ### What breaks it **The decision prompt carries the entire loop.** `getAgentDecisionPrompt` is the only thing standing between a targeted follow-up and a wasted iteration. Getting consistent JSON and sound decisions out of it takes real testing and a few-shot example set, which this post omits. Everything else here is plumbing you can read; this is the part you have to tune. **Malformed JSON costs an iteration.** The fallback in `evaluateContext` catches a parse failure and returns `NEEDS_MORE_INFO` with a generic refinement query. That's the safe default, and it's also a full search-plus-evaluate cycle spent on a formatting error rather than a knowledge gap. With `MAX_RAG_ITERATIONS` at 3, one bad parse is a third of your budget. **Nothing here shows it beats static RAG.** The honest state of this is that I haven't measured it. Demonstrating that self-correction wins would mean RAGAS, human evaluation, or fact-checking metrics on the same queries against a single-shot baseline, and until that exists the argument for the loop is structural rather than empirical. **Everything upstream can fail mid-loop.** The embedding service can go down on iteration two, with an iteration's worth of context already accumulated. Degrade to what you have rather than throwing the whole run away. --- ### When agentic retrieval is the wrong choice The loop turns one embed plus one search plus one LLM call into up to three of each, plus synthesis. That's latency and money, and there are cases where it buys nothing. - **The first retrieval is reliably sufficient.** Lookup-shaped queries against a corpus that answers them in one hop. The evaluation call is pure overhead here, and it can only ever tell you `COMPLETE`. - **You're on a tight latency budget.** Each iteration is a sequential embed, search, and model call. Three iterations plus synthesis is a serial chain, and nothing in this design overlaps them. - **The corpus doesn't contain the missing piece.** The loop can only re-query the same vector store. If the gap the critique names isn't in the index, better follow-up queries just spend iterations finding that out more expensively. - **You can't test the decision prompt.** Without a way to check that follow-up queries are actually targeting the named gap, you've added a non-deterministic step to your control flow and no way to tell when it misfires. - **You have no evaluation harness.** If you can't compare against a static baseline, you can't tell whether the extra calls improved the answer or just made the run longer. --- ### The shift The interesting part of building this wasn't better answers. It was watching the model stop being a text generator and start being a control plane: running a workflow, deciding the next step, adapting as it went. That reframe has a price attached, and it's the one worth internalising. Every iteration is another call, another wait, another charge. The competitive programming habit is the right one here, where every token and every call counts, because in a loop those costs multiply instead of adding. - **Give the model a way to say what's missing.** A status alone can't drive a search; a critique can. - **Bound the loop from outside the model.** Iteration caps, token caps, and `k` are yours, not its. - **Split the model by role.** Cheap and fast for the decision, capable for the synthesis. The next thing this wants is specialised agents chosen by the *type* of gap found, a fact-checker for one, a summariser for another, a deep-dive for a third, which is where the MoE analogy stops being an analogy. > A system that knows what it doesn't know yet is worth more than one that answers faster. This loop is the smallest version of that I could build without a framework. --- Source: https://himanshuat.com/blogs/self-correcting-rag-agentic-recursive-retrieval ======================================================================== --- title: "Serverless GPUs for bursty training" description: "Our finetuning load arrived in bursts, a few hard days then a quiet week, and a rented GPU box covers that badly. The split we ended up running, and the one number I still can't give you." date: "June 17, 2025" url: "https://himanshuat.com/blogs/serverless-gpus-for-bursty-training" --- # Serverless GPUs for bursty training For most of a month our GPU sat idle. Then a client batch landed, we wanted three LoRA variants trained by the weekend, and one node wasn't enough. The bill was never what pushed us off the default. On a team small enough that nobody's job is to look after infrastructure, the box became somebody's job. > A machine you never rebuild is a machine whose configuration nobody knows. That cost arrives as a training run failing for a reason nobody can reconstruct, not as a line on the invoice. ### Chapter 0 — What a burst is A burst is work with two properties: it's separated by idle, and the only thing that matters about it is wall clock to finish. Nothing happens for days. Then 3,000 new occlusion images land and a finetune kicks off, or a batch of client renders has to be through the pipeline before a review call. That's a different animal from the sustained work running through the same hardware: reading papers, poking at a ComfyUI graph, running the pipeline against a handful of images over and over to see whether a change to the mask dilation did anything. That work is interactive, it wants a machine that's already warm, and it runs for hours at low utilisation. Size one node for the burst and it's idle the rest of the month. Size it for the interactive work and the burst takes as long as it takes. --- ### Step 1 — Separate the two workloads before you size anything The two kinds of work want opposite things, and averaging them gives you a node that serves neither. | | sustained R&D | burst jobs | |---|---|---| | shape | hours at low utilisation | days idle, then everything at once | | what matters | time to first result | wall clock to finish | | tolerates a cold start | no | yes, if the job is long | | where we run it | dedicated EC2 node | Modal | So: dedicated Amazon EC2 nodes for sustained training and R&D, Modal for the bursts in both training and inference, and S3 in the middle holding source images, generated outputs and logs so neither side owns the data. ```mermaid flowchart LR W[work] --> Q{sustained or burst} Q -->|sustained| E[EC2 node] Q -->|burst| M[Modal fan-out] E --> S[(S3)] M --> S ``` The dedicated node is where you keep a shell open. The serverless functions are where you run things you want four of. *We also ran diffusion inference on ShaktiCloud, an India-hosted GPU cloud, for work that needed to stay closer to home. That's a separate decision and it didn't change the burst story.* Cost is the argument everyone reaches for and it's the weakest one, because a reserved node is cheap per hour and it's easy to convince yourself the idle time is fine. Two other things did more damage. The box accumulates state. Somebody installs a build of xformers to test something, somebody else pins a different diffusers version, a checkpoint gets dropped in `/home` because the volume was full, and three weeks later a training run fails and nobody can reconstruct why. And a single box serialises the burst. Comparing a rank 32 adapter against the same data at a different learning rate is two runs, and on one node they're two runs back to back. The whole point of a LoRA at rank 32 is that it's small and cheap enough to train several of, and that advantage disappears if the hardware makes them sequential. **The tell:** if your GPU's last month of utilisation is a flat low line with three spikes in it, you are paying for the spikes all month and still queueing behind them when they arrive. --- ### Step 2 — Define the image in code, next to the function The part of Modal that changed how I work isn't the GPU. It's that the container image is defined in Python next to the function that runs in it. No Dockerfile drifting out of sync with what's actually installed on the box, and no ambiguity about what a run was executed against. ```python # modal_app.py """Modal app for AuraX burst workloads: LoRA training and batch inference. The image below is the whole environment. If a run reproduces here, it reproduces anywhere, which is the property the long-lived EC2 node had stopped having. """ import modal app = modal.App("aurax-burst") CUDA_TAG = "12.4.1-devel-ubuntu22.04" image = ( modal.Image.from_registry(f"nvidia/cuda:{CUDA_TAG}", add_python="3.11") .apt_install("git", "libgl1", "libglib2.0-0") .pip_install( "torch==2.4.0", "diffusers==0.30.0", "transformers==4.44.0", "accelerate==0.33.0", "safetensors==0.4.4", "bitsandbytes==0.43.3", # AdamW8bit "peft==0.12.0", "boto3==1.34.162", "pillow==10.4.0", ) # Kohya-ss is pinned to a commit, not a branch. A training script that moves # under you is the same failure as a box whose packages nobody tracks. .run_commands( "git clone https://github.com/kohya-ss/sd-scripts /opt/sd-scripts", "cd /opt/sd-scripts && git checkout 5b6c2e2 && pip install -r requirements.txt", ) .env({"HF_HOME": "/weights/hf", "PYTHONUNBUFFERED": "1"}) ) # Weights are a volume, not a layer. See below for why. weights = modal.Volume.from_name("aurax-weights", create_if_missing=True) outputs = modal.Volume.from_name("aurax-outputs", create_if_missing=True) secrets = [modal.Secret.from_name("aws-s3"), modal.Secret.from_name("hf-token")] ``` Pinning Kohya-ss to commit `5b6c2e2` rather than a branch is the single line in that file I'd defend hardest. Training scripts in this ecosystem move fast and quietly, and a LoRA you can't retrain is a LoRA you can't fix. **The tell:** if you can't read off the repo which version of the training script a given run executed against, the image isn't the deployment yet. --- ### Step 3 — Put the weights in a volume, not in a layer The rule I settled on: a layer holds anything that changes when the code changes, a mount holds anything that changes on its own schedule. | what it is | changes when | where it goes | |---|---|---| | system packages, dependencies, training script | the code changes | image layer | | datasets, generated outputs, logs | on their own schedule | volume mount | | model weights | rarely, but every output moves with them | volume, versioned by path | Model weights sit across that line, which is why they're the awkward case. They behave like data, being large and opaque and rarely edited. They behave like code, because swapping the checkpoint changes every output, so a run is only reproducible if you know which one it used. Bake them and you buy reproducibility by rebuilding gigabytes for an unrelated dependency bump. Mount them and you buy fast rebuilds at the cost of a mutable input, unless you version by path and never overwrite in place, which is what we ended up doing. The first version of this baked the checkpoint into the image. It worked, and it was wrong. Every dependency bump rebuilt and re-uploaded gigabytes for no reason, and worse, the image became the thing that determined which checkpoint a run used, which is exactly the coupling you were trying to avoid. ```python # modal_app.py (continued) @app.function( image=image, volumes={"/weights": weights}, secrets=secrets, timeout=60 * 60, ) def sync_weights(keys: list[str]) -> None: """Pull base checkpoints and adapters from S3 into the shared volume. Run this once when a checkpoint changes. Every training or inference function then mounts /weights read-mostly and starts against whatever is there, so the checkpoint in use is a property of the volume rather than of the image. """ import os from pathlib import Path import boto3 s3 = boto3.client("s3") bucket = os.environ["AURAX_WEIGHTS_BUCKET"] for key in keys: dest = Path("/weights") / key if dest.exists(): print(f"skip {key} (present, {dest.stat().st_size / 1e9:.1f} GB)") continue dest.parent.mkdir(parents=True, exist_ok=True) print(f"pull {key}") s3.download_file(bucket, key, str(dest)) # Commit makes the new files visible to other containers mounting this volume. weights.commit() print("volume committed") ``` There's a subtlety here that cost me an afternoon. A volume mounted into a running container is a snapshot, so a function that started before `commit()` will not see the new file. If you kick off training in the same session that uploaded the checkpoint, order it explicitly rather than assuming. **The tell:** if bumping a pip pin re-uploads gigabytes of checkpoint, your weights are in the wrong place. --- ### Step 4 — Fan out the runs you would otherwise queue With the image and the weights sorted, the training function is thin. It shells out to the same Kohya-ss invocation we'd run on the dedicated node, with the same rank, alpha, learning rate and optimiser, so a run is comparable across the two environments. ```python # train_burst.py """Fan out LoRA training runs. One container per run, all of them at once.""" import modal from modal_app import app, image, outputs, secrets, weights @app.function( image=image, gpu="L40S", volumes={"/weights": weights, "/out": outputs}, secrets=secrets, timeout=60 * 60 * 6, retries=modal.Retries(max_retries=1, backoff_coefficient=1.0), ) def train_lora(run: dict) -> str: """Train one adapter and return the path of the artifact it produced. `run` carries only the things we vary between experiments. Everything else is fixed on purpose, so that a diff between two runs is a diff in the data or in one hyperparameter, never in the environment. """ import subprocess name = run["name"] cmd = [ "accelerate", "launch", "--num_cpu_threads_per_process", "8", "/opt/sd-scripts/sdxl_train_network.py", "--pretrained_model_name_or_path", "/weights/flux/base.safetensors", "--train_data_dir", f"/weights/datasets/{run['dataset']}", "--output_dir", "/out/lora", "--output_name", name, "--resolution", "1024,1024", "--network_module", "networks.lora", "--network_dim", "32", "--network_alpha", "16", "--learning_rate", str(run.get("lr", 1e-4)), "--lr_scheduler", "cosine", "--train_batch_size", "4", "--gradient_accumulation_steps", "4", "--gradient_checkpointing", "--optimizer_type", "AdamW8bit", "--mixed_precision", "bf16", "--save_every_n_epochs", "1", ] subprocess.run(cmd, check=True) outputs.commit() artifact = f"/out/lora/{name}.safetensors" print(f"finished {name} -> {artifact}") return artifact @app.local_entrypoint() def main(): runs = [ {"name": "drape_lr1e4", "dataset": "drape_5k", "lr": 1e-4}, {"name": "drape_lr5e5", "dataset": "drape_5k", "lr": 5e-5}, {"name": "occ_lr1e4", "dataset": "occlusion_3k", "lr": 1e-4}, {"name": "occ_lr5e5", "dataset": "occlusion_3k", "lr": 5e-5}, ] for artifact in train_lora.map(runs): print("ready:", artifact) ``` `train_lora.map(runs)` is the whole argument in one line. Four runs, four containers, four GPUs, one wall clock. On a single node that's four runs in sequence plus a decision about which one you care about enough to run first. I stopped making that decision. That's worth more to a research loop than it sounds, because the runs you skip are disproportionately the ones that would have surprised you. **The tell:** if you are ranking experiments by which one to run first, and the ranking exists only because the hardware makes them sequential, that ranking is what fan-out deletes. --- ### Step 5 — Amortise the load inside the container Every burst container pays for pulling the image, mounting the volume and loading a large checkpoint into VRAM before it does any work you care about. For a six hour training run that overhead is noise. For a batch of a few dozen renders it is not. Three knobs mattered, and only one of them is free. Keeping containers warm between calls during an active session helps a lot and costs you idle time you're paying for, so it's a session-level decision rather than a permanent setting. Loading the checkpoint once per container and processing many images inside that container beats one image per invocation by a wide margin. And keeping the checkpoint on a volume rather than pulling it from S3 per container removes the largest single chunk of that startup, which is the practical reason the volume design in Step 3 exists. ```python # batch_infer.py """Batch inference: amortise the load over many images inside one container.""" import modal from modal_app import app, image, outputs, secrets, weights @app.cls( image=image, gpu="L40S", volumes={"/weights": weights, "/out": outputs}, secrets=secrets, scaledown_window=300, # keep the container warm during an active batch max_containers=8, # ceiling, not a target ) class TryOnWorker: @modal.enter() def load(self): """Runs once per container, before any request is served.""" import torch from diffusers import FluxFillPipeline self.pipe = FluxFillPipeline.from_pretrained( "/weights/flux/fill", torch_dtype=torch.bfloat16 ).to("cuda") self.pipe.load_lora_weights("/weights/lora/aurax_merged.safetensors") print("pipeline resident") @modal.method() def render(self, job: dict) -> str: out = self.pipe( prompt=job["prompt"], image=job["image"], mask_image=job["mask"], num_inference_steps=30, height=1024, width=1024, ).images[0] path = f"/out/renders/{job['id']}.png" out.save(path) return path ``` The `@modal.enter()` split is the whole trick. Model loading happens once per container lifetime, request handling happens per call. Getting that boundary wrong is the most common way a serverless GPU setup ends up looking slow and expensive, and it looks exactly like a cold start problem when it's really a code structure problem. **The tell:** if your per-image cost doesn't fall as the batch gets longer, the model is loading per request. --- ### What breaks it **Cold starts, and I can't tell you where the crossover is.** There's a job length below which you should just run the thing on the node that's already warm. I would like to tell you where that sits for us. I can't, and I'd rather say so than make one up. We never instrumented cold starts properly, we tuned by feel, and the only latency figure I trust from that whole period is a lab measurement of roughly **12 seconds** for a full-body try-on on an RTX 4090, single image, nothing to do with production serving. Every number in the code above is a configuration value we set, not something we measured. If you're evaluating this for your own work, measure it yourself against your checkpoint size, because that's the variable that dominates and it's the one I can't hand you. **Latency that presents as a capacity problem.** You change one line in a sampler config and run it to see what happened. The job is correct, the batch completes, nothing is slow in any sense a dashboard would report. But the gap between pressing enter and seeing a pixel is long enough that you tab away, and by the time the render lands you've lost the thread of why you made the change. So you stop asking one question at a time and start batching them, which is a worse way to work, because the value of an interactive loop is that each result changes what you ask next. I spent longer than I'd like to admit throwing parallelism at it. **Debugging through a keyhole.** When a training run dies on a box you own, you SSH in and look. The process is still there, the half-written checkpoint is still on disk, `nvidia-smi` tells you what the memory actually did, and you can open a shell against the same environment and poke at the tensor that came out wrong. When a burst function dies, the container is gone. What you have is the logs you had the foresight to print, and if you didn't print the thing you now need, you get to run the job again to find out. That changes how you write the code more than how you debug it. You print more and print earlier, on the theory that a log line you don't need costs nothing while one you didn't write costs a re-run. And you care much more about whether a failure reproduces locally. So anything we didn't already understand stayed on the dedicated node until we did. --- ### When serverless GPUs are the wrong choice The split above is a good default for our shape of work. It is overhead for several others. - **The work is interactive.** Reading papers, tuning a ComfyUI graph, iterating on mask dilation against a handful of images. That work is intolerant of a wait between pressing enter and seeing a result, and a warm box you already own is the correct answer. - **The job is short relative to startup.** Every container pays for the image pull, the volume mount and the checkpoint load before it does anything useful. Below some job length that overhead dominates, and I don't know where that line is for you any more than I know it for us. - **You don't understand the failure modes yet.** Serverless is a good place to run a job whose failure modes you can predict, and a bad place to find out what a new one is. The container is gone before you can look at it. - **You never want four of something.** Fan-out is most of the value here. If your experiments genuinely run one at a time, a single node is simpler and you keep the shell. --- ### The shift The thing I'd change first is embarrassingly cheap. I'd instrument the startup path on day one, even crudely: a timestamp at container entry, one after the checkpoint loads, one at first output, written to the same S3 bucket as everything else. It costs almost nothing, and it's the difference between this post having an answer in it and this post admitting it doesn't. The other thing I'd change is treating the two environments as one system earlier. Once the training invocation was byte-identical between the node and the burst function, a whole category of "it works on the box but not on Modal" disappeared. That identity should have been the first thing I built rather than something I converged on. - **Split by workload shape, not by cost.** Interactive work wants a warm box; burst work wants many cold ones. - **Version by path, rebuild by code.** Weights on a volume, everything the behaviour depends on in a layer. - **Instrument the startup path before you need the number.** You will want it exactly once, and by then the period you'd have measured is over. The next piece of this is what to do with four adapters once you've trained them in parallel, which turns out to be a much harder problem than getting the GPUs. --- Source: https://himanshuat.com/blogs/serverless-gpus-for-bursty-training ======================================================================== --- title: "StyleX- Replacement of TailWindCSS?" description: "A look at StyleX, Facebook's CSS-in-JS library, and how it compares to Tailwind." date: "December 16, 2023" url: "https://himanshuat.com/blogs/stylex-new-frontend-styling-framework" --- # StyleX- Replacement of TailWindCSS? ## Table of contents 1. [What is StyleX?](#what-is-stylex) 2. [Atomic CSS](#atomic-css) 3. [Dynamic styles](#dynamic-styles) 4. [Code splitting](#code-splitting) 5. [Dead code elimination](#dead-code-elimination) 6. [Type safety](#type-safety) 7. [Using StyleX](#using-stylex) - [Example 1: creating styles](#example-1-creating-styles) - [Example 2: dynamic styles](#example-2-dynamic-styles) - [Example 3: theming](#example-3-theming) 8. [StyleX vs. TailwindCSS](#stylex-vs-tailwindcss) - [Syntax and API](#syntax-and-api) - [Performance and bundle size](#performance-and-bundle-size) - [Features and trade-offs](#features-and-trade-offs) - [Ecosystem and community](#ecosystem-and-community) 9. [Conclusion](#conclusion) ## What is StyleX? StyleX is a CSS-in-JS library from Facebook, aimed at the problems that older CSS-in-JS approaches run into on large codebases. A few things stand out. ## Atomic CSS StyleX uses atomic CSS: it generates small, reusable classes you compose together, which keeps styling modular and easier to maintain as a project grows. ## Dynamic styles You can drive styles from props, state, or context with a declarative syntax, so styles adapt at runtime without much ceremony. ## Code splitting StyleX splits CSS into chunks loaded on demand, which trims the initial payload and speeds up first load. ## Dead code elimination At build time StyleX strips styles you don't use, so only what the app actually references ends up in the final bundle. ## Type safety StyleX works with TypeScript and Flow, so you get autocompletion and type checking on style properties and values. ## Using StyleX ### Example 1: creating styles ```jsx import { stylex } from "stylex"; const styles = stylex.create({ container: { padding: "10px", backgroundColor: "lightblue", }, }); function MyComponent() { return
Hello StyleX!
; } ``` ### Example 2: dynamic styles ```jsx import { stylex } from "stylex"; function DynamicComponent({ isActive }) { const dynamicStyles = stylex.use({ color: isActive ? "green" : "red", }); return
Dynamic Style
; } ``` ### Example 3: theming ```jsx import { stylex, ThemeProvider, useTheme } from "stylex"; const themes = { light: { background: "white", text: "black" }, dark: { background: "black", text: "white" }, }; function ThemedComponent() { const theme = useTheme(); const styles = stylex.create({ container: { backgroundColor: theme.background, color: theme.text, }, }); return
Themed Content
; } function App() { return ( ); } ``` ## StyleX vs. TailwindCSS Here's how StyleX stacks up against Tailwind, which is what a lot of people would be switching from. ### Syntax and API StyleX has its own syntax and API, which is a real shift if you're coming from Tailwind's utility classes. There's a learning curve. ### Performance and bundle size Its atomic styling and code splitting are meant to give it an edge over Tailwind on bundle size and runtime cost. ### Features and trade-offs Both have their strengths. StyleX leans into atomic and dynamic styles; the cost is learning a new model, which tends to pay off more on large projects than on small ones. ### Ecosystem and community Tailwind has a large ecosystem and community. StyleX is newer and smaller, though that may change as it opens up. ## Conclusion StyleX is a credible alternative to Tailwind, especially on large projects, and having Facebook behind it doesn't hurt. Whether it displaces Tailwind is another question, but it's worth a look. To try it, check the official website and the GitHub repository. --- Source: https://himanshuat.com/blogs/stylex-new-frontend-styling-framework ======================================================================== --- title: "Generic aesthetic scorers hate e-commerce photography" description: "We pointed an open aesthetic scorer at a batch of generated product shots and it ranked the moody, low-key ones highest. That's a defensible opinion about photographs and the wrong answer for a catalogue. This is the scorer we built instead, and the parts of it I can't put a number on." date: "April 29, 2025" url: "https://himanshuat.com/blogs/teaching-a-model-what-a-catalogue-looks-like" --- # Generic aesthetic scorers hate e-commerce photography The first time we ran an off-the-shelf aesthetic scorer over a batch of generated product images, it put the dramatic ones on top. Deep shadows, warm rim light, a face half in darkness, a background you couldn't read. Genuinely nice images. Completely unusable as catalogue assets. > The scorer wasn't making a mistake about photographs. It was answering a different question from the one we were asking, and its answers were legible enough that you can go a long way before you notice the mismatch. We were optimising against somebody else's definition of good. ### Chapter 0 — What a scorer is for Two jobs, and they have different tolerance for error. **Filtering the training corpus.** After the mechanical pass (resolution, sharpness, aspect ratio) there's a harder question: is this the kind of image we want the model to produce? A perfectly sharp, well-exposed garment shot on a beach at golden hour passes every technical check and is the wrong training example for a catalogue model. The aesthetic scorer makes that call at a scale where a human can't. **Ranking generation candidates.** At inference you generate more than one candidate and you have to pick one. This is where the scorer earns its keep most obviously, because the alternative is showing every candidate to a person, and that doesn't scale past a demo. Those two uses interact, and it took me a while to see how. ```mermaid flowchart LR S[scorer] --> F[filter corpus] F --> M[train model] M --> C[candidates] C --> S S --> P[picked] ``` If the scorer filters the corpus, the model learns to produce images the scorer likes. If the scorer then ranks that model's candidates, it's grading work its own preferences shaped. Closed loops narrow. Whatever the scorer slightly over-rewards gets amplified twice, once through what the model was trained on and once through what gets selected at generation time. We kept a holdout the scorer never touched, and ran human spot checks on candidate ranking rather than trusting top-1 blindly. *That's a mitigation, not a solution.* --- ### Step 1 — Write down the question your scorer answers Aesthetic scorers get trained on human ratings of art and photography. That's the data that exists, and the people rating it are answering "is this a good photograph". A good photograph tends to have mood. The scorer learns that correlation honestly. It isn't broken. A catalogue product shot wants close to the opposite. | | good photograph | catalogue asset | |---|---|---| | lighting | low-key, strong contrast | even, so no shadow reads as a stain | | background | shallow depth of field, unreadable | neutral, so the page design can hold it | | colour | graded | corrected to match the physical garment | | framing | crop and occlusion are mood | hem, sleeve, and neckline all visible | Colour is the one with money behind it. **The return rate is the metric that eventually pays for all of this**, and a garment that arrives a different colour than the page showed comes back. Every property in the right column is something the general scorer treats as flat, boring, or under-directed, including vignette-free lighting. So you get a scorer that is confidently wrong in a consistent direction. Not noise, which would be easy to spot. A systematic pull toward darker and moodier, and if you use it as a training filter you build that pull into everything downstream. Build the scorer around a question a human can answer in two seconds. "Does this belong in this brand's catalogue" is that question. "Is this aesthetic" is not, and a scorer trained on the second will quietly answer a third one you never asked. **The tell:** if the images your scorer ranks highest are ones you would never put on a product page, it is answering somebody else's question well. --- ### Step 2 — Collect labels in pairs, never on a scale The labelling side is deliberately dull. Pairwise comparison rather than an absolute rating, because a human asked to score an image from one to ten will drift across a session, anchor on whatever they saw last, and use a different part of the scale after lunch. Asked which of two images is better for this brand's catalogue, the same person is far more consistent. ```python # scorer/collect_preferences.py """Pairwise preference collection for the brand-centric aesthetic model. Raters see two images and answer one question: which of these belongs in this brand's catalogue? Absolute 1-to-10 scoring was tried first and abandoned; the same rater would not reproduce their own scores across sessions, and two raters would not agree on what a 7 meant. """ import json import random from dataclasses import asdict, dataclass from datetime import datetime, timezone from pathlib import Path @dataclass class Comparison: brand_id: str left: str right: str winner: str # "left", "right", or "skip" rater_id: str flipped: bool # whether display order was reversed recorded_at: str def sample_pair(pool: list[str], seen: set[frozenset]) -> tuple[str, str] | None: """Draw an unseen pair. Returns None once the pool is exhausted.""" for _ in range(200): left, right = random.sample(pool, 2) if frozenset((left, right)) not in seen: return left, right return None def present(left: str, right: str) -> tuple[str, str, bool]: """Randomise display order so position bias averages out.""" flipped = random.random() < 0.5 return (right, left, True) if flipped else (left, right, False) def record(out_path: Path, comparison: Comparison) -> None: """Append one comparison. Skips are kept: they mark ambiguous pairs.""" with out_path.open("a") as handle: handle.write(json.dumps(asdict(comparison)) + "\n") def session(brand_id: str, rater_id: str, pool: list[str], out_path: Path, target: int = 100) -> None: seen: set[frozenset] = set() collected = 0 while collected < target: pair = sample_pair(pool, seen) if pair is None: break left, right = pair shown_a, shown_b, flipped = present(left, right) choice = prompt_rater(shown_a, shown_b) # UI layer, elided seen.add(frozenset(pair)) record(out_path, Comparison( brand_id=brand_id, left=left, right=right, winner=choice, rater_id=rater_id, flipped=flipped, recorded_at=datetime.now(timezone.utc).isoformat(), )) collected += 1 print(f"brand={brand_id} rater={rater_id} collected={collected} " f"pairs_exhausted={pair is None}") ``` Two details in there carry more weight than they look like they do. Display order is randomised so position bias averages out. And skips are stored rather than dropped, because a pair a rater refuses to call is information: usually two images differing along an axis the brand doesn't care about, and knowing which axes those are tells you **what the scorer should be blind to.** Pairwise beats absolute scoring for anything a human labels by eye, and the gap is larger than I expected. **The tell:** if a rater can't reproduce their own scores from last week, the scale is the problem, not the rater. --- ### Step 3 — Make the score a routing decision, not a verdict What sits behind the scoring call itself is not something I'm writing up here. The interface is the part that matters for anyone building this, and the interface is small. ```python # scorer/rank_candidates.py """Rank generation candidates with the brand-centric scorer. The scorer is a black box at this boundary on purpose. Everything below depends only on score(image, brand_id) -> float. """ from dataclasses import dataclass from typing import Protocol import yaml class BrandScorer(Protocol): def score(self, image_path: str, brand_id: str) -> float: """Higher is more catalogue-appropriate for this brand.""" ... @dataclass(frozen=True) class BrandPolicy: brand_id: str accept_threshold: float # below this, do not surface the candidate review_threshold: float # between the two, route to human review max_surfaced: int def load_policy(path: str, brand_id: str) -> BrandPolicy: with open(path) as handle: config = yaml.safe_load(handle)[brand_id] return BrandPolicy(brand_id=brand_id, **config) def rank(scorer: BrandScorer, candidates: list[str], policy: BrandPolicy) -> dict[str, list[str]]: """Split candidates into surfaced, review, and rejected buckets.""" scored = sorted( ((path, scorer.score(path, policy.brand_id)) for path in candidates), key=lambda item: item[1], reverse=True, ) surfaced, review, rejected = [], [], [] for path, value in scored: if value >= policy.accept_threshold and len(surfaced) < policy.max_surfaced: surfaced.append(path) elif value >= policy.review_threshold: review.append(path) else: rejected.append(path) print(f"brand={policy.brand_id} candidates={len(candidates)} " f"surfaced={len(surfaced)} review={len(review)} rejected={len(rejected)}") return {"surfaced": surfaced, "review": review, "rejected": rejected} ``` **The three-bucket split rather than a straight top-N is the part I'd keep in any version of this.** A single threshold forces the scorer to be right, and it isn't reliably right near the boundary. The middle band is where the human effort goes, and it stays affordable because most candidates fall clearly on one side or the other. The thresholds live in per-brand config rather than in code, and I'm deliberately not printing values here. They aren't comparable across brands. The scorer is calibrated separately for each one, so the same number means a different thing for two clients, and publishing a pair of them would imply a shared scale that doesn't exist. They also moved often enough that any value I quoted would be a snapshot of one week. **The tell:** if your pipeline has one threshold, you have asked a model to be right at the exact point where it is least reliable. --- ### Step 4 — Calibrate per brand, because "clean" is not one thing Here's the part that doesn't resolve. One brand means a white cyclorama, hard even light, a bright high-key look, and skin that's been smoothed. Another means daylight through a window, visible shadow under the jaw, warm neutrals, and skin with texture left in it. Both will tell you, in the same words, that they want clean commercial imagery. An image one of them signs off on will get rejected by the other. So the scorer has to be calibrated per brand, or at minimum per brand profile. **The labelling cost doesn't amortise the way you'd hope: every new client is partly a new labelling problem**, and that cost is the real price of this approach, paid again at each onboarding. Individual raters aren't stable either. Order effects are real. Fatigue is real. Somebody who has looked at four hundred images in a row starts rewarding novelty, which is exactly wrong for catalogue work. **The tell:** if two clients use the same word for what they want and your scorer serves both from one calibration, one of them is about to reject a batch you were confident in. --- ### Step 5 — Say plainly what you haven't measured I'm not going to put numbers on any of this. No rater counts, no agreement statistics, no accuracy figures for the scorer. We ran this as a working process rather than as a study. Quoting a number would imply a level of rigour the setup didn't have, and I'd rather say that than publish a figure that sounds authoritative and isn't. If you're building the same thing and you want to know whether your raters agree, you'll have to run that properly yourself. It's worth doing. Saying which parts you haven't measured is more useful to a reader than an invented number, and it keeps you honest with yourself about which of your components you actually trust. **The tell:** if you can't name the study that produced a number, don't publish the number. --- ### What breaks it **Reward hacking, and it arrives quietly.** If a scorer rewards even lighting and neutral backgrounds, and you select generations against it, the model drifts toward flat and safe. Every individual image passes. The catalogue as a whole gets duller, and nobody can point at the image where it went wrong. The only defence I found was keeping images the scorer never selected in front of a human periodically. **Staleness.** A brand's art direction changes seasonally, and a scorer trained on last season's approved assets will start rejecting the new direction with total confidence. The refresh path has to exist from the beginning, not arrive as a later feature. **The score standing in for correctness.** The scorer answers whether an image fits a brand's catalogue. It does not answer whether the try-on is correct. Garment fidelity, occlusion, and drape are measured separately. Letting an aesthetic score cover correctness means a beautifully lit image of the wrong garment ranked first, and that is the single worst output this system can produce. --- ### When training your own scorer is the wrong choice - **You're at demo scale.** If a person can look at every candidate, that person is the scorer and they're better at it. The entire argument for a model here is that human review doesn't survive past a demo. - **Your question really is "is this a good photograph".** The open scorers answer that one honestly and for free. They only point the wrong way for catalogue work, because catalogue work wants the opposite of mood. - **The brand hasn't settled what clean means.** The labels are the asset, and you cannot collect labels against a target that's still moving. Wait for the art direction, then label. - **What you actually need is correctness.** Fidelity, occlusion, and drape need their own measurement. An aesthetic scorer will rank the wrong garment first and look confident doing it. --- ### The shift The mistake wasn't picking a bad scorer. It was picking a scorer without writing down the question, and then treating its output as a verdict instead of a routing decision. - **Score a question a human can answer in two seconds.** "Does this belong in this brand's catalogue", not "is this aesthetic". - **Ask for comparisons, not scores.** A rater who can't reproduce a 7 can reliably pick a winner. - **Route, don't rule.** Three buckets with a person in the middle band absorbs most of the error the model is going to make. The concrete next step is a per-brand calibration set collected at onboarding: a short pairwise session run by the brand's own art director before we generate anything for them, so the scorer starts from their definition of clean instead of ours. It costs the client part of an afternoon. > It also removes the argument that otherwise happens three weeks later, over images you've already delivered. --- Source: https://himanshuat.com/blogs/teaching-a-model-what-a-catalogue-looks-like ======================================================================== --- title: "The backtest is the experiment, and most of them aren't controlled" description: "Self-taught quant research mostly fails on bias-laden backtests rather than on the math. So I front-loaded the pitfall literature before running anything, and it changed what I built first — infrastructure, not strategies." date: "August 12, 2026" url: "https://himanshuat.com/blogs/the-backtest-is-the-experiment" --- # The backtest is the experiment, and most of them aren't controlled The failure mode I was warned about isn't running out of math. It's producing a beautiful equity curve that means nothing, believing it, and finding out eighteen months later. Most self-taught quant research fails on bias-laden backtests and untargeted reading. Not on stochastic calculus. > A backtest is an experiment. Almost nobody treats it like one, and an uncontrolled experiment returns a number whether or not there's an effect. So I inverted the usual order. Before the first backtest, I read the literature on how backtests lie. Before the first strategy, I built the machinery that makes a result reproducible. --- ### Chapter 0 — What a backtest actually claims A backtest claims: *had I run this rule over this period, I would have earned this.* That claim is only as good as the counterfactual. Every bias is a way the counterfactual is false while the arithmetic is correct. The named ones I had to internalise before I could trust anything: - **Lookahead.** The rule used information that didn't exist at decision time. - **Survivorship.** The universe silently excludes what died. - **Selection.** You chose the period, the universe, or the parameters after seeing results. - **P-hacking.** You ran many specifications and reported the one that worked. - **Regime change.** The relationship existed and then stopped. None of these produce an error. They all produce a *number*, which is what makes them dangerous. --- ### Step 1 — Read the pitfall canon before the textbooks The reading order matters more than the reading list. Pitfall literature first, because it changes how you read everything after it. The four I front-loaded: - **Bailey & López de Prado (2014), *The Deflated Sharpe Ratio*.** How to discount a Sharpe ratio for the number of trials that produced it. - **Bailey, Borwein, López de Prado & Zhu, *Pseudo-Mathematics and Financial Charlatanism*.** The case that backtest overfitting is the field's default state rather than its edge case. - **Harvey, Liu & Zhu (2016), *…and the Cross-Section of Expected Returns*.** Multiple testing applied to the factor zoo — what happens to published factors when you correct for how many were tried. - **Aronson, *Evidence-Based Technical Analysis*.** Bias enumeration done properly. The shared argument across all four is one idea: **a result's credibility depends on how many results you didn't report.** A Sharpe ratio without a trial count is an unlabelled measurement. Reading these after building a strategy would have been useless. I'd have had a result to defend. **The tell:** if you can't say how many specifications you tried before the one you're showing, you can't interpret its Sharpe. --- ### Step 2 — Build reproducibility before you build signal This is the part I'd have skipped and shouldn't have. It reads like infrastructure procrastination. It's actually the control group. What has to exist before the first real backtest: - **Experiment tracking** — mlflow or wandb, logging *every* backtest with its config. Not the good ones. Every one. This is what makes the trial count in Step 1 a number instead of a guess. - **Data versioning** — dvc, or hashed parquet snapshots. A result you can't re-run against the same bytes isn't a result. - **Deterministic notebooks** — seeded RNG, library versions pinned in a lockfile. - **Walk-forward and purged k-fold utilities** — mlfinlab implements López de Prado's versions. Standard k-fold leaks across time-series folds; purging is what removes the leak. The ordering argument is the same one from Step 1, turned into engineering. If tracking is opt-in, you'll log the runs you're proud of, and your trial count will be wrong in the direction that flatters you. **The tell:** if your experiment tracker has fewer runs in it than you've actually run, its purpose has already been defeated. --- ### Step 3 — Replicate before you originate The learning loop isn't reading papers. It's re-implementing them. The three-pass protocol I use, and the honest cost of each: | pass | time | what you get | |---|---|---| | 1 | ~5 min | title, abstract, intro, headers, conclusion. Decide if pass 2 is worth it. | | 2 | ~1 hr | figures, tables, methodology. Skip proofs. Note data, universe, sample period, claimed result. | | 3 | ~3 hr | re-derive the math, re-implement the core experiment, sanity-check the result. | The target that matters: **at least five pass-3 papers a year.** Pass-1 everything else. Five is a small number and it still sounds ambitious once you've done one. That gap — between how easy a paper is to read and how hard it is to reproduce — is most of what pass 3 teaches. The foundational set to replicate against is well-defined: Markowitz (1952), Sharpe (1964), Black & Scholes (1973), Fama & French (1993) and (2015), Carhart (1997), Jegadeesh & Titman (1993), Asness/Moskowitz/Pedersen, Hou/Xue/Zhang, and Gu, Kelly & Xiu (2020). **The tell:** if your reproduction matches the paper on the first try, check your data for lookahead before you celebrate. --- ### Step 4 — Pick one niche and let it constrain the reading Untargeted reading is the second named failure mode, and it's the one that feels productive while it's happening. The fix is structural: pick a narrow niche first, and let it decide what's worth a pass 3. Without an anchor, every interesting paper is equally interesting, which means the queue grows faster than it drains. The infrastructure that supports this is unglamorous — an SSRN account following FEN and CMBO feeds, the arXiv q-fin daily for q-fin.ST, q-fin.PR and q-fin.CP, a Zotero queue, and one paper a week with notes. **The tell:** if your "papers to read" queue is growing monotonically, you don't have a niche yet. --- ### What breaks it **Tracking that's easy to bypass.** The moment logging a run takes an extra step, exploratory runs stop being logged, and exploratory runs are exactly the ones the deflated Sharpe correction needs to count. **Purging that isn't applied consistently.** Purged cross-validation on the final model and plain k-fold during exploration means the leak happened during the part where you made the decisions. **Replication that stops at "the number matches."** A reproduction can agree with the paper and still be wrong in the same way the paper was. Matching a published result is evidence about your implementation, not about the effect. **Reading the canon and then not applying it.** Knowing about selection bias does not remove selection bias. The corrections have to be mechanical — in the tracker, in the CV splitter — or they don't happen under deadline. --- ### When this much rigour is the wrong choice - **You're learning the mechanics of an instrument.** Pricing a vanilla option by hand to understand it doesn't need experiment tracking. The machinery is for claims about edge, not for exercises. - **The result will never inform a decision.** A study run purely to understand a paper's method can stop at pass 3. Reproducibility infrastructure earns its cost when someone might act on the output. - **You haven't picked a niche.** Building walk-forward tooling before you know what you're testing is the same procrastination as untargeted reading, wearing engineering clothes. - **The effect size is enormous and mechanical.** Some relationships are accounting identities rather than discoveries. Multiple-testing corrections aren't the binding constraint there. The rigour scales with how much the claim would cost you if it were false. For most of what I'll produce in the first year, that cost is my own time — which is exactly the budget the failure mode consumes. --- ### The shift The instinct is to learn the math, then build a strategy, then check it. That order puts verification last, which is where it always gets cut. Front-loading the pitfall canon inverted it. Now the reading list is chosen by a niche, the trial count is a logged number rather than a memory, and the cross-validation splitter refuses to leak whether or not I remember to ask it to. - **Read how backtests lie before you run one.** Afterwards, you have a result to defend. - **Make the corrections mechanical.** Knowing about p-hacking doesn't prevent p-hacking; a tracker that logs every run does. - **Count reproductions, not papers.** Five pass-3s a year beats a hundred abstracts. > A Sharpe ratio is a measurement. The number of times you measured is part of the measurement, and it's the part that never makes it into the chart. --- Source: https://himanshuat.com/blogs/the-backtest-is-the-experiment ======================================================================== --- title: "554 million ticks and the wrong reason mmap loses" description: "Everyone repeats that mmap is bad for large scans. On Apple Silicon that turned out to be true, false, and true again depending on working-set size — and the mechanism everybody names for it is not the mechanism doing the work." date: "August 12, 2026" url: "https://himanshuat.com/blogs/the-fault-policy-not-the-page-size" --- # 554 million ticks and the wrong reason mmap loses I had a 22 GB tick CSV, 554 million rows, and two ways to read it. The received wisdom said use `read()`. Crotty et al. argued at CIDR'22 that mmap is a poor choice for database scans, and the paper has been repeated often enough that it functions as a rule. Then I measured it, and mmap won by 1.2 to 1.3x. Then I made the file bigger and mmap lost by a factor of 0.78. > The rule isn't wrong. It's unconditional, and the thing it's conditional on is working-set size relative to RAM. The other half of the received wisdom is the explanation: Apple Silicon uses 16 KiB base pages, so of course the fault behaviour differs from 4 KiB x86. That part turns out to be the wrong mechanism entirely. --- ### Chapter 0 — What the two strategies actually are **Buffered `read()`** copies bytes from the page cache into a buffer you own. You reuse one small buffer forever. The kernel reads ahead in large chunks because it knows you're streaming. **mmap** maps the file into your address space and lets the MMU fault pages in on demand. No copy, no buffer management. You touch a byte, the hardware traps, the kernel populates the page. The pitch for mmap is that it deletes the copy. The case against it is that it converts a sequential I/O problem into a page-fault problem, and page faults are a bad way to express "read the next 64 KiB." Both claims are true. Which one dominates is an empirical question, and it's the one the literature settled at the level of mechanism rather than measurement. --- ### Step 1 — Hold the parser constant, or measure nothing The trap in any I/O benchmark is that you end up measuring your parser. The parser here is a hand-rolled, validating, zero-copy byte scanner, identical across every arm. Isolated from I/O it sustains around **11 GB/s** — comfortably faster than the disk. That number is the whole license for the rest of the experiment. If the parser is never the bottleneck, whatever separates the arms is I/O. The arms: | arm | what it tests | |---|---| | buffered `read()` | the baseline the literature recommends | | `mmap` | demand paging, no advice | | `mmap` + `MADV_SEQUENTIAL` | can advice close the gap? | | `mmap` + `MADV_WILLNEED` | can eager population close the gap? | | parallel `mmap` stream | does concurrency change the shape? | | `csv` + `serde` | a stand-in for what most people actually write | Each arm runs in a **fresh subprocess**, so `getrusage` peak RSS and the minor/major fault counters are clean rather than accumulated. Warmup runs are discarded. The working-set sweep uses newline-aligned byte limits so every arm parses the same rows. **The tell:** if your fast arm and your slow arm don't share a parser, you have a parser benchmark with an I/O label on it. --- ### Step 2 — Sweep the working set until the answer changes sign A single file size gives you a single verdict, which is how a conditional result gets published as an unconditional one. Sweeping produces a clean crossover: | working set vs RAM | mmap relative to `read()` | |---|---| | fits | **1.2–1.3x faster** | | exceeds | **0.78x — slower** | The crossover falls at roughly **one-third to two-thirds of physical memory.** Not at 100%. The strategies diverge well before you exhaust RAM. > A benchmark that reports one number for "mmap vs read" has picked a side of a crossover and called it a finding. **The tell:** if you never varied the input size, you don't know which side of the crossover your result is on. --- ### Step 3 — Attribute the slowdown per fault, not per second Wall-clock tells you mmap lost. It doesn't tell you why, and "mmap is slow for scans" is a restatement, not a cause. The fault counters do tell you: - Buffered `read()` takes **essentially zero disk faults at any size.** The kernel reads ahead; nothing traps. - mmap's disk faults climb to **1.34 million**, one 16 KiB page at a time. That's the mechanism. Not copy avoidance, not TLB pressure, not cache behaviour. The mapped arm is paying a trap per page and the buffered arm is paying none. ```mermaid flowchart TD F["22 GB tick CSV, 554M rows"] --> R["buffered read()"] F --> M["mmap"] R --> RB["small reused buffer,
kernel prefetches large
sequential reads"] RB --> RF["~0 major faults
at any working-set size"] M --> MF["one fault per
16 KiB page"] MF --> WS{"working set
vs RAM"} WS -->|"fits"| MIN["minor faults,
page already resident
1.2-1.3x faster than read()"] WS -->|"exceeds"| MAJ["major faults climb to 1.34M,
each a blocking single-page read
0.78x slower than read()"] RF --> OUT["read() wins out-of-core"] MAJ --> OUT MIN --> OUT2["mmap wins in-RAM"] ``` Why the two strategies diverge once the working set exceeds RAM: the buffered arm never reaches the decision node at all. **The tell:** a performance claim you can't attribute to a counter is a story. Find the counter that moves. --- ### Step 4 — Check whether your mechanism survives crossing the OS This is the step that changed the paper. The obvious explanation for the fault count is page size: macOS uses 16 KiB pages, Linux uses 4 KiB, so Linux should fault **more**, not less. Four times more, if page size is what's driving it. It faults roughly **four times less.** Two things on the Linux side account for it. Fault-around populates **64 KiB per fault** rather than one page. And anonymous memory gets folded into huge pages. So the same scan, with pages one quarter the size, takes a quarter of the faults — because the OS decided to populate more per trap. The control that closes it: **x86-64 and AArch64 Linux agree with each other.** Same fault-population behaviour across two instruction sets. > The effect is the operating system's fault-population policy. Not the page size, and not the ISA. That inverts the intuition the field has been carrying. The page-size story predicts the opposite sign of the effect that actually shows up. **The tell:** if your explanation predicts a direction, go find a platform where the direction should flip. If it doesn't flip, the explanation is wrong. --- ### Step 5 — Try to rescue the losing arm before you write it off `madvise` is the obvious rescue. If the problem is fault frequency, tell the kernel what's coming. It doesn't work on macOS: - **`MADV_SEQUENTIAL`** buys about **15%** — and leaves the fault count *unchanged*. It's making each fault cheaper, not making fewer of them. The mechanism I identified in Step 3 is untouched. - **`MADV_WILLNEED`** *doubles* major faults and is the **slowest arm in the entire experiment.** The eager-population hint makes it actively worse. The 15% is the more interesting of the two, because it's the number that would have let me claim advice "helps" if I'd only looked at wall-clock. The fault counter is what says it's not addressing the cause. **The tell:** when a fix improves time but not the counter you blamed, you fixed something else. --- ### What breaks it **Timing under virtualization is not timing.** The cross-OS comparison runs under virtualized and emulated Linux. So I take **only fault counts** from it, which are architectural, and never timing. This constrains the claim: I can say Linux's fault policy populates more per trap, and I cannot say Linux is faster. **One machine, one disk, one schema.** All measurements are on a single Apple Silicon machine with one NVMe device and one dataset schema. The crossover ratio is a property of that configuration until a second storage device says otherwise. **The parser being fast is load-bearing.** At 11 GB/s the parser is far from the bottleneck. On a slower parser — a `csv` + `serde` pipeline, which is what most people actually deploy — the I/O difference compresses, because the bottleneck moves. The result is about I/O strategy for a workload where I/O is the constraint. **Bare-metal Linux confirmation and a second storage device are the outstanding experiments** before this is submittable. Targeting arXiv (cs.OS, cs.PF, cs.DB), then a workshop such as DaMoN or HotStorage. --- ### When mmap is the wrong choice - **Your working set exceeds RAM.** This is the case the literature was right about. Above the crossover you pay 1.34 million traps for a copy you weren't spending much on. - **You're streaming once and discarding.** Demand paging's advantage is reuse. A single sequential pass over data you'll never touch again is exactly the shape buffered reads with readahead were built for. - **Your parser is the bottleneck.** If you're at `csv` + `serde` speeds rather than 11 GB/s, the I/O strategy is not what's costing you, and swapping it is effort spent on the wrong layer. - **You need predictable latency.** Faults are a distribution, not a constant. `read()` taking essentially zero disk faults at any size is a stability property, not just a speed one. Conversely: below the crossover, on a re-scanned working set, with a parser fast enough to notice, mmap's 1.2–1.3x is real and worth taking. --- ### The shift I went in expecting to confirm a known result and got a conditional one, with the wrong mechanism attached to it. The measured claim is narrower and more useful than the rule it replaces. mmap doesn't lose to `read()`. mmap loses *above a working-set threshold that lands between one-third and two-thirds of RAM*, and the thing that sets that threshold is how many bytes your OS populates per fault. - **Sweep the input size.** A single-size benchmark picks a side of a crossover and calls it a law. - **Attribute to a counter.** Wall-clock says who won. Fault counts say why, and "why" is what transfers to your workload. - **Test the mechanism where it should flip.** Page size predicted Linux would fault more. It faulted four times less, and that's how the real cause surfaced. > The received wisdom had the right verdict for the wrong reason, which meant it was also right in cases where it shouldn't have been and wrong in cases nobody checked. --- Source: https://himanshuat.com/blogs/the-fault-policy-not-the-page-size ======================================================================== --- title: "Three clouds, one inference path" description: "Inference ran on three providers because one brand's images could not leave the country, one workload wanted a machine you could leave dirty, and the other sat at zero for four days and then arrived as a catalogue drop. This is the routing rule that made it survivable, and the bill I never measured." date: "September 17, 2025" url: "https://himanshuat.com/blogs/three-clouds-one-inference-path" --- # Three clouds, one inference path A graph that ran fine on one host produced a black image on another. Same weights, same manifest, a different driver and CUDA combination underneath. That is where the genuinely annoying evenings went, and it is the recurring bill for running inference in three places at once. > Three providers was never a cost decision. It was a decision about which promises we could make to a brand, and I still can't put a number on what it cost us. At some point in 2025 our GPU work was spread across dedicated Amazon EC2 nodes, Modal.com, and ShaktiCloud, an India-hosted GPU cloud. Whenever I drew that on a whiteboard, the first question back was some version of "why not pick one". It's a fair question and the answer isn't a spreadsheet. We never built the measurement harness that would let me put a cost comparison in front of you, so this post is the reasoning, and I'll say where the reasoning ran ahead of the evidence. ### Chapter 0 — What a pool is A **pool** is where a unit of GPU work can be scheduled. Not a vendor, not a region: a place with a shape. We had three. Long-lived EC2 nodes for training and R&D. Serverless GPU on Modal that scales with the queue. In-region capacity on ShaktiCloud for diffusion inference under a residency constraint. The rest of this post is about how a job gets assigned to one of them, and what it cost to keep three of them alive. --- ### Step 1 — Split the work by shape, not by vendor Training a LoRA expert is a long, stateful, messy job. You want a machine you can leave dirty: datasets on local disk, a Kohya-ss config you're editing between runs, checkpoints accumulating, a tmux session you reattach to the next morning. Rank 32 at alpha 16 on a few thousand curated images is not something you re-provision from scratch every time you change one number. R&D on the graph itself has the same shape. Half that work is watching an intermediate mask in a browser and deciding the dilation radius is wrong. Inference is a different animal. A brand doesn't send a steady trickle of garments. They send nothing for four days and then a catalogue drop, because someone finished a shoot and wants the whole set through the pipeline this afternoon. Between bursts the useful amount of GPU capacity is close to zero. Holding dedicated nodes for that pattern means paying for idle silicon most of the week and still being short when the drop lands. | workload | shape | what it needs | pool | |---|---|---|---| | LoRA training, eval | long, stateful, checkpoint-heavy | a machine you can leave dirty | EC2 dedicated | | graph R&D | interactive, half of it is looking at masks | same machine, reattached | EC2 dedicated | | catalogue inference | zero for days, then a burst | capacity that appears and leaves | Modal burst | If it had stopped there, this would be a two-line post about baseline versus burst. **The tell:** if a job needs a machine you can leave dirty between runs, serverless will fight you. If a job's useful capacity is zero for four days at a stretch, dedicated will bill you for those four days. --- ### Step 2 — Treat residency as a constraint, not a preference The thing that made it three was not technical. Indian fashion brands care where unreleased campaign imagery sits. Think about what we were handed: a model image and a garment image for a collection that hasn't launched, from a brand whose commercial advantage that quarter is that nobody has seen it yet. In conversations with larger apparel groups, the questions we got were never about FID. They were about where the file goes, whose hardware it touches, and whether the output could be produced without leaving the country. ShaktiCloud is an India-hosted GPU cloud, and we used it for diffusion inference. That was the reason. Not price, not a benchmark we ran against the alternatives, but the ability to say that a given brand's inference happens in-country and to have that be true rather than a claim we'd walk back under scrutiny. Being concrete about what the requirement forces matters here, because "data residency" sounds like a checkbox and behaves like an architecture. The weights are the easy part. The merged model and the LoRA experts are ours, they contain no customer pixels, and we can copy them anywhere. What's constrained is what the brand hands us and what we hand back. S3 was holding three classes of object under one label: the input images, the generated outputs, and the logs. The third is the one teams forget, because a diffusion pipeline's logs accumulate image keys, mask previews, and sometimes a downscaled render of whatever went wrong. > If the constraint covers the picture, it covers the thumbnail of the picture sitting in your error tracker. So a constrained tenant needs its own bucket in the right place, the worker has to be told which one rather than knowing a default, and I couldn't pull a failing image down to my laptop to look at it. There was a second, quieter benefit. GPU availability in-region is genuinely uneven, and having a provider whose capacity isn't competing with every other workload on the planet meant that when we needed machines, sometimes we could get machines. *I can't quantify that for you. It's an operational impression from months of trying to get GPUs, not a study.* **The tell:** if you can still pull a constrained brand's failing image down to your laptop, the residency claim is a slide, not an architecture. --- ### Step 3 — Make the placement call in exactly one place Three providers is survivable only if the decision about which one runs a job is made once, explicitly, and never leaks into the worker code. The queue from our API already gave us the seam: workers pull jobs, so placement is just a routing rule about which pool a job goes to. The order of the checks is the whole design. ```mermaid flowchart TD J[job] --> K{kind} K -->|train or eval| E[EC2 dedicated] K -->|inference| R{residency IN} R -->|yes| S[Shakti in-region] R -->|no| M[Modal burst] ``` ```python # infra/placement.py from dataclasses import dataclass from enum import Enum class Pool(str, Enum): """Where a unit of GPU work can be scheduled.""" EC2_DEDICATED = "ec2_dedicated" # long-lived nodes, training and R&D MODAL_BURST = "modal_burst" # serverless GPU, scales with the queue SHAKTI_IN_REGION = "shakti_in_region" # India-hosted, diffusion inference @dataclass class Tenant: tenant_id: str residency_required: str | None # e.g. "IN", or None if unconstrained @dataclass class Job: job_id: str kind: str # "train" | "inference" | "eval" tenant: Tenant manifest_id: str priority: str = "standard" # "standard" | "interactive" def place(job: Job, queue_depth: int) -> Pool: """Pick the pool for a job. Order matters. Residency is a constraint, not a preference, so it is checked before anything about capacity or cost. Everything after it is a judgement call we can revisit without breaking a promise to a brand. """ if job.kind in {"train", "eval"}: # Long, stateful, checkpoint-heavy. Never worth a cold container. return Pool.EC2_DEDICATED if job.tenant.residency_required == "IN": return Pool.SHAKTI_IN_REGION if job.priority == "interactive" and queue_depth == 0: # A human is waiting on this one and there is warm capacity. return Pool.MODAL_BURST return Pool.MODAL_BURST def explain(job: Job, queue_depth: int) -> str: """Human-readable reason, written into the job record for auditing.""" pool = place(job, queue_depth) if pool is Pool.EC2_DEDICATED: reason = "training or evaluation workload" elif pool is Pool.SHAKTI_IN_REGION: reason = f"residency constraint {job.tenant.residency_required}" else: reason = "burst inference" return f"{job.job_id} -> {pool.value} ({reason})" if __name__ == "__main__": abfg = Tenant("abfg", residency_required="IN") partner = Tenant("api_partner", residency_required=None) print(explain(Job("job_a1", "inference", abfg, "vton-2025-09-a"), queue_depth=12)) print(explain(Job("job_b2", "inference", partner, "vton-2025-09-a"), queue_depth=0)) print(explain(Job("job_c3", "train", partner, "drape-r32"), queue_depth=0)) ``` The part I care about in that file is `explain`. Every job record carried the pool it ran on and the reason it went there. When a brand asked us where their images had been processed, the answer was a database query rather than an engineer's recollection. That single habit paid for itself more than once. Notice also what isn't in there. No cost term, no bidding on spot capacity, no attempt to be clever. > A placement policy that optimises something you aren't measuring is a policy that will surprise you. **The tell:** when a brand asks where their images were processed, you should be running a query, not searching your memory. --- ### Step 4 — Give every artifact one key layout and a provenance sidecar The other half of making this tolerable was refusing to let storage fragment. Everything, from every pool, wrote to S3 with the same key layout, and nothing in the worker was allowed to invent a path. ```python # infra/artifacts.py import json import io from datetime import datetime import boto3 s3 = boto3.client("s3") BUCKET = "aurax-artifacts" def key_for(tenant_id: str, job_id: str, kind: str, ext: str) -> str: """Build the one canonical key for an artifact. Layout is tenant first so that a per-brand export or deletion is a prefix operation, and date second so lifecycle rules can expire intermediates without touching delivered outputs. """ day = datetime.utcnow().strftime("%Y/%m/%d") return f"{tenant_id}/{day}/{job_id}/{kind}.{ext}" def put_output(tenant_id: str, job_id: str, image_bytes: bytes, meta: dict) -> dict: """Write a generated image plus its provenance sidecar.""" image_key = key_for(tenant_id, job_id, "output", "png") meta_key = key_for(tenant_id, job_id, "provenance", "json") s3.upload_fileobj(io.BytesIO(image_bytes), BUCKET, image_key) s3.put_object( Bucket=BUCKET, Key=meta_key, Body=json.dumps(meta, indent=2).encode(), ContentType="application/json", ) return {"image": image_key, "provenance": meta_key} if __name__ == "__main__": provenance = { "job_id": "job_a1", "pool": "shakti_in_region", "manifest_id": "vton-2025-09-a", "seed": 2847193044, "lambda_drape": 0.6, "lambda_occlusion": 0.4, "steps_primary": 30, "steps_refine": 10, } print(json.dumps(provenance, indent=2)) print(key_for("abfg", "job_a1", "output", "png")) ``` A provenance sidecar next to every image sounds like bookkeeping until the week you need it. Ours recorded the pool, the manifest, and the seed, which meant any output could be traced back to a specific set of weights on a specific kind of machine. When you're running the same graph in three places, that file is the only thing standing between you and a guessing game. **The tell:** pick a delivered image at random. If you can't name the pool, the manifest, and the seed behind it, you're running three pipelines rather than one. --- ### Step 5 — Keep the provider surface down to glue The worker itself was built to be uninteresting. One container image. Model weights synced from object storage on start rather than baked in. No provider SDK imported anywhere in the inference path, and no assumption about the local filesystem beyond a scratch directory. ```bash #!/usr/bin/env bash # infra/bootstrap_worker.sh # Prepare a GPU host to run the try-on graph. Identical on every provider. set -euo pipefail MANIFEST_ID="${1:?usage: bootstrap_worker.sh }" SCRATCH="${SCRATCH_DIR:-/scratch}" MODELS="${SCRATCH}/models" mkdir -p "${MODELS}" "${SCRATCH}/comfy" # 1. Pull the manifest that defines this worker's behaviour. aws s3 cp "s3://aurax-artifacts/manifests/${MANIFEST_ID}.json" "${SCRATCH}/manifest.json" # 2. Fetch exactly the weights it names, verifying each hash. python3 - "${SCRATCH}/manifest.json" "${MODELS}" <<'PY' import hashlib, json, pathlib, subprocess, sys manifest = json.load(open(sys.argv[1])) dest = pathlib.Path(sys.argv[2]) for model in manifest["models"]: target = dest / model["file"] if not target.exists(): subprocess.check_call([ "aws", "s3", "cp", f"s3://aurax-artifacts/weights/{model['file']}", str(target), ]) digest = hashlib.sha256(target.read_bytes()).hexdigest() print(f"{model['role']:<12} {model['file']:<32} {digest[:8]}") PY # 3. Check out the custom nodes at the commits the manifest pins. python3 infra/sync_nodes.py "${SCRATCH}/manifest.json" "${SCRATCH}/comfy" # 4. Hand over to the queue consumer. No HTTP server, no provider specifics. exec python3 -m worker.consume --manifest "${SCRATCH}/manifest.json" ``` Every provider difference that survived after this lived in about thirty lines of glue per provider: how you get credentials, how you request a machine, how logs come back. That was the target, and mostly we hit it. **The tell:** if moving a workload to a new provider is a project rather than a config change, the provider has leaked into the worker. --- ### Four things to do before you add the second provider Write the residency requirement into the tenant record before you need it, not after a brand asks. We retrofitted that field, and retrofitting a constraint is always worse than starting with it. Build the accounting harness at the same time as the second pool. The moment you have two pools you have a comparison question, and the data to answer it has to be collected from the beginning. You cannot reconstruct it later from invoices. My inability to show numbers in this post traces back to skipping that step for one more sprint, repeatedly. Keep the provider surface tiny and treat it as an asset. The bootstrap script above is unglamorous, and it is the reason moving a workload was a config change. Be suspicious of your own justifications. Two of our three pools had a constraint behind them that I can still defend today. The third was defensible when we set it up and became partly habit afterwards, and I didn't revisit it as often as I should have. --- ### What breaks it **Three of everything, forever.** Three sets of credentials and three permission models. Three ways a machine can fail to appear, with three different error messages meaning "no capacity right now". Three driver and CUDA combinations, which is where the black-image evenings came from. And three mental models to keep loaded, which for a team our size is the expensive part even though it never shows up anywhere you can point at. **There is no single dashboard, and building one keeps losing.** The concrete version is the question "why hasn't this image come back yet", which we got often enough that I can still recite the sequence. Check the job record in Postgres for status and pool. If it's queued, check Redis for depth and whether any worker is claiming. If it's running, go to that pool's console, which is a different console with a different login and a different idea of what a log line looks like. Building the unified view meant normalising three log formats, so we kept not doing it. **Your own signal gets noisier.** Splitting inference across pools splits your traffic, which means every operational instinct you build is built on partial data. Warm capacity behaves differently when a pool handles most of your jobs than when it handles a slice. By dividing the work we degraded our own ability to read it. --- ### When three pools is the wrong choice **Nobody has ever asked you where the file goes.** Then the third pool is pure cost and evenings. Residency is what justified ours, and without that question arriving from a real account, you are paying the three-of-everything bill for nothing. **You wanted redundancy.** These pools were not backups for each other. A residency-constrained job can't fail over to a pool outside the country, which is the point of the constraint. The dedicated EC2 nodes were sized for training runs we'd already committed to, so pointing burst inference at them during an outage would have cost us the capacity we were holding them for. Redundancy is a different design, with spare capacity we didn't have. **One vendor can give you all three properties.** In-country capacity, scale-to-zero for the bursty half, and machines you can hold for a week. That wasn't on offer. If it existed we'd have moved the same afternoon, and I'd recommend you do too. **You need to defend a per-image cost.** I can't, because we never ran the comparison properly. We had invoices and impressions. If someone tells you they know the cost-per-image of a diffusion pipeline across three providers, ask them how they attributed idle time on the dedicated nodes. We couldn't have answered that. **You're small and the constrained accounts aren't the ones that matter.** A team our size probably should have picked one provider and lived with it: serve the brands we could serve, and tell the ones with a residency requirement we'd be ready later. I went the other way because those questions were coming from exactly the accounts that would have made the company work, so simplicity would have meant choosing to stay small on purpose. *I still think the call was right, and I hold it less confidently than I did in January.* --- ### The shift We chose placement on constraints we could state (this brand's images stay in India, this training run needs a machine that persists) rather than on a per-image figure we could defend. That was the honest position at the time. It is also the reason this post has code in it and no numbers. - **Check residency before capacity.** A constraint you break to save money is a promise you shouldn't have made. - **Log the reason, not just the destination.** `explain` turned a recurring customer question into a query. - **Build the accounting harness with the second pool, not after the third.** You cannot reconstruct it from invoices. The concrete next step, if the work continues, is to make the placement policy read utilisation instead of hard-coding it: have the router consult live queue depth and warm-capacity signals per pool, and log what it would have chosen against what it did choose. Run that in shadow mode for a month and you finally have the dataset that makes the cost conversation possible instead of theoretical. > Everything I can't tell you in this post is the price of never having measured the rest. --- Source: https://himanshuat.com/blogs/three-clouds-one-inference-path ======================================================================== --- title: "Part 1: The Attention Mechanism - Building the Core of the Transformer" description: "An RNN has to compress the whole input sequence into one fixed-size vector before it is allowed to predict anything. Attention drops that constraint. This post builds it from raw tensor operations — projections, scaled dot product, softmax, weighted sum — in NumPy first, then PyTorch." date: "May 10, 2024" url: "https://himanshuat.com/blogs/understanding-transformers-part-1-attention-from-scratch" --- # Part 1: The Attention Mechanism - Building the Core of the Transformer An RNN folds the entire input sequence into a single fixed-size vector, and only then is it allowed to make a prediction. Summarize a novel in one sentence, then guess its genre from the sentence. That is the bottleneck. The longer the sequence, the more the vector drops, and nothing in the architecture tells you which part went missing. Attention removes the constraint. Rather than squashing everything into one vector, the model weighs the importance of different parts of the input as it processes each element. > Attention is a primitive for selective information processing. Language models are just the place it showed up first. I want to build an architecture that gets somewhere near the flexibility of the human brain, out of interpretable components, eventually using Mixture-of-Experts (MoE) to handle specialized functions. That is a long way off, and it starts with the primitives. This series starts with the one I find most interesting: the attention mechanism. No framework abstracting it away here — raw tensor operations and a clear look at how the model focuses. ### Chapter 0 — Query, key, value The core components are **Query**, **Key**, and **Value**. The clearest way in is a database lookup. | component | what it is | the question it answers | |---|---|---| | Query (Q) | your search term | what am I looking for? | | Key (K) | the label attached to each item | what do I have that might match? | | Value (V) | the item itself | if there's a match, what comes back? | In self-attention, every token in the sequence plays all three roles. When the model processes one token, that token's query is compared against *every* token's key to score relevance. The scores then weight a sum over the values, producing a new representation of the query token that carries its context. The shape of the whole computation is a fan-out and a merge. One embedding tensor becomes three, two of those meet to form the scores, and the third rejoins at the end. ```mermaid flowchart LR E[embeddings] --> Q[Q] E --> K[K] E --> V[V] Q --> S[scores] K --> S S --> SC[scale] SC --> SM[softmax] SM --> O[weighted sum] V --> O ``` Every arrow in that diagram is a matrix multiplication, which is what makes it fast on a GPU. There are no iterative loops anywhere in it. We start in NumPy, where the operations are fully visible, then move to PyTorch. The inputs are token embeddings: dense vectors standing in for words. ```python import numpy as np import torch import torch.nn as nn import math # --- Hyperparameters --- seq_len = 5 # Example: "The quick brown fox jumps" d_model = 8 # Embedding dimension for each token d_k = d_model # Often d_k = d_v = d_model / num_heads, but for simplicity here, d_k = d_model d_v = d_model # The dimension of the Value vectors # --- 1. Synthetic Input Embeddings --- # batch_size=1 for now, but attention scales perfectly for batches. # Shape: (batch_size, seq_len, d_model) input_embeddings = np.random.rand(1, seq_len, d_model) print(f"Input Embeddings Shape: {input_embeddings.shape}\n") ``` `d_k = d_model` is a teaching simplification. In a real Transformer, `d_k` is `d_model` divided by the number of heads, which is what Part 2 is about. --- ### Step 1 — Project the embeddings into three spaces Project the input embeddings into Query, Key, and Value space using three separate weight matrices. Three matrices, not one, because the model needs to learn a different transformation for each role. ```python # --- Weights for Linear Projections --- # These would be learned during training. For now, random initialization. # Shape: (d_model, d_k) for Q and K, (d_model, d_v) for V W_q = np.random.rand(d_model, d_k) W_k = np.random.rand(d_model, d_k) W_v = np.random.rand(d_model, d_v) # --- Compute Q, K, V Matrices --- # (batch_size, seq_len, d_model) @ (d_model, d_k) -> (batch_size, seq_len, d_k) Q = input_embeddings @ W_q K = input_embeddings @ W_k V = input_embeddings @ W_v print(f"Q Matrix Shape: {Q.shape}") print(f"K Matrix Shape: {K.shape}") print(f"V Matrix Shape: {V.shape}\n") ``` Each row of `Q`, `K`, and `V` is one token's query, key, or value vector. These are transformed representations, not copies, which is what lets the model attend differently depending on which role a token is playing. **The tell:** three weight matrices, three printed shapes. If `Q`, `K`, and `V` are the same tensor, nothing downstream can tell the three roles apart. --- ### Step 2 — Score every query against every key Take the dot product between the Query matrix and the transpose of the Key matrix. For each token's query vector this measures similarity against every other token's key vector. Higher dot product, higher similarity. ```python # --- Attention Scores --- # (batch_size, seq_len, d_k) @ (batch_size, d_k, seq_len) -> (batch_size, seq_len, seq_len) # Using `transpose(1, 2)` to swap the last two dimensions for K. attention_scores = Q @ K.transpose(0, 2, 1) # K.T in NumPy often means (..., K.T), need explicit permute. # For a single batch, it's simpler: Q[0] @ K[0].T # For full batching: np.einsum('bsd,bTd->bsT', Q, K) is often more robust, but @ works too. print(f"Attention Scores Shape: {attention_scores.shape}") print(f"Sample Attention Scores (first token's scores):\n{attention_scores[0, 0, :]}\n") ``` The result has shape `(batch_size, seq_len, seq_len)`. For a given batch and query token `i`, `attention_scores[batch_idx, i, j]` is how much token `i` attends to token `j`. Note the transpose. `K.T` on a batched NumPy array does not do what you want here; you need an explicit permute of the last two axes. **The tell:** the score matrix comes out square in the sequence dimension. Any other shape and you transposed the wrong axes. --- ### Step 3 — Scale by the square root of `d_k` Dot products grow large when `d_k` is high, and large scores push softmax into a region where the gradient is tiny and learning stalls. So scale the scores by $\frac{1}{\sqrt{d_k}}$. It is the easiest line in the formula to skim past and it is the one holding training together. ```python # --- Scaling --- scaled_attention_scores = attention_scores / np.sqrt(d_k) print(f"Sample Scaled Attention Scores:\n{scaled_attention_scores[0, 0, :]}\n") ``` **The tell:** the divide sits before the softmax. Scale afterwards and the row no longer sums to 1, which the check in Step 4 will catch for you. --- ### Step 4 — Softmax the rows into a distribution Apply softmax row-wise, over the last dimension. This converts the scaled scores into probability distributions, so the attention weights for each query sum to 1. Now `attention_weights[batch_idx, i, j]` really is the probability that token `i` focuses on token `j`. ```python # --- Softmax Function --- def softmax(x, axis=-1): e_x = np.exp(x - np.max(x, axis=axis, keepdims=True)) # Subtract max for numerical stability return e_x / e_x.sum(axis=axis, keepdims=True) attention_weights = softmax(scaled_attention_scores, axis=-1) print(f"Attention Weights Shape: {attention_weights.shape}") print(f"Sample Attention Weights (first token's):\n{attention_weights[0, 0, :]}") print(f"Sum of weights for first token: {np.sum(attention_weights[0, 0, :]):.4f}\n") ``` The max subtraction inside `softmax` is not decoration. Exponentiating raw scores is where numerical stability goes. **The tell:** print the row sum. If it is not approximately 1.0, you softmaxed over the wrong axis. --- ### Step 5 — Take the weighted sum of the values Multiply the attention weights by the Value matrix. Each query token's new representation is a weighted sum of every Value vector in the sequence, with the attention probabilities as the weights. Tokens that scored high contribute more. ```python # --- Weighted Sum of Values --- # (batch_size, seq_len, seq_len) @ (batch_size, seq_len, d_v) -> (batch_size, seq_len, d_v) output = attention_weights @ V print(f"Output Shape: {output.shape}\n") print("Self-Attention Mechanism Complete (NumPy)!\n") ``` `output` holds the contextually enriched representation of each token. Every token now knows about the relevant parts of its surrounding context. The shapes are the whole story of the pipeline, so it is worth reading them in one place: | stage | shape | |---|---| | input embeddings | `(batch_size, seq_len, d_model)` | | Q, K | `(batch_size, seq_len, d_k)` | | V | `(batch_size, seq_len, d_v)` | | scores, scaled scores, weights | `(batch_size, seq_len, seq_len)` | | output | `(batch_size, seq_len, d_v)` | **The tell:** the output has the same shape as `V`. Same number of rows in, same number out; only the contents changed. --- ### Step 6 — Port it to PyTorch and assert the two agree NumPy is good for seeing the operations. PyTorch is what you would actually run on a GPU. The operations are identical, expressed on tensors, and the assert at the bottom is the point of writing it twice. ```python # --- PyTorch Implementation --- print("--- PyTorch Implementation ---") # Convert NumPy arrays to PyTorch tensors input_embeddings_torch = torch.tensor(input_embeddings, dtype=torch.float32) # Linear layers (equivalent to W_q, W_k, W_v) # We use nn.Linear for standard PyTorch layers. # Note: In a real Transformer, these would be `nn.Linear(d_model, d_k)` etc. # For simplicity, we'll manually create weight tensors matching our NumPy example, # but usually, you'd instantiate nn.Linear layers. linear_q = nn.Linear(d_model, d_k, bias=False) linear_k = nn.Linear(d_model, d_k, bias=False) linear_v = nn.Linear(d_model, d_v, bias=False) # Initialize weights to match our NumPy random ones for consistency (optional) with torch.no_grad(): linear_q.weight.copy_(torch.tensor(W_q.T, dtype=torch.float32)) # PyTorch linear expects (out_features, in_features) linear_k.weight.copy_(torch.tensor(W_k.T, dtype=torch.float32)) linear_v.weight.copy_(torch.tensor(W_v.T, dtype=torch.float32)) # 1. Compute Q, K, V Q_torch = linear_q(input_embeddings_torch) K_torch = linear_k(input_embeddings_torch) V_torch = linear_v(input_embeddings_torch) print(f"Q_torch Shape: {Q_torch.shape}") print(f"K_torch Shape: {K_torch.shape}") print(f"V_torch Shape: {V_torch.shape}\n") # 2. Calculate Attention Scores # Q @ K.transpose(-2, -1) performs (batch, seq_len, d_k) @ (batch, d_k, seq_len) attention_scores_torch = torch.matmul(Q_torch, K_torch.transpose(-2, -1)) print(f"Attention Scores_torch Shape: {attention_scores_torch.shape}\n") # 3. Scaling scaled_attention_scores_torch = attention_scores_torch / math.sqrt(d_k) # 4. Softmax attention_weights_torch = torch.softmax(scaled_attention_scores_torch, dim=-1) print(f"Attention Weights_torch Shape: {attention_weights_torch.shape}") print(f"Sum of weights for first token (PyTorch): {torch.sum(attention_weights_torch[0, 0, :]):.4f}\n") # 5. Weighted Sum of Values output_torch = torch.matmul(attention_weights_torch, V_torch) print(f"Output_torch Shape: {output_torch.shape}\n") print("Self-Attention Mechanism Complete (PyTorch)!") # Verify results are close between NumPy and PyTorch assert np.allclose(output, output_torch.numpy(), atol=1e-6) print("NumPy and PyTorch outputs are consistent!\n") ``` Two implementations of the same math, checked against each other, is a cheap way to find out which one you got wrong. **The tell:** `np.allclose(output, output_torch.numpy(), atol=1e-6)` passes. When it does not, start at the weight transpose — `nn.Linear` stores its weight as `(out_features, in_features)`, which is why `W_q.T` goes in and not `W_q`. --- ### What breaks it **Skipping the scale.** With a high `d_k`, unscaled dot products get large, softmax saturates, gradients go small, and training stalls without ever raising an error. Nothing in the shapes tells you this happened. The loss curve does, eventually. **Softmax without the max subtraction.** `np.exp` on raw scores is where numerical stability dies. Subtracting the row max before exponentiating costs one line and removes the problem. **Transpose conventions, in both libraries.** In NumPy, `K.T` on a batched array is not the transpose you want; you need an explicit permute of the last two axes, or `np.einsum('bsd,bTd->bsT', Q, K)`. In PyTorch, `nn.Linear` expects `(out_features, in_features)`, so the NumPy weight goes in transposed. Both mistakes produce tensors of plausible shape that compute the wrong thing. --- ### When single-head attention from scratch is the wrong choice The version built above is a teaching configuration. It is the right thing to write once and the wrong thing to ship. - **You need more than one relation at a time.** Here `d_k = d_model` and there is exactly one head, so the model gets exactly one way of relating tokens. Multi-head attention splits `d_model` across heads to get several, which is the whole reason Part 2 exists. - **You want a model that runs, not a model you can read.** NumPy is here because the operations are visible in it. If the goal is training on a GPU, that path is a detour — go to PyTorch and stay there. - **The sequence is long.** The score matrix is `(batch_size, seq_len, seq_len)`: every token compared against every token. That shape is fixed by the mechanism, and it grows with the square of the sequence length whether or not the extra comparisons carry information. - **The mechanism is not what you are working on.** Hand-rolling the projections and the softmax buys understanding of the internals. Understanding of the internals is the only thing it buys. If you are debugging a training run rather than learning the primitive, the framework's implementation is the one to use. --- ### The shift Building attention from raw tensor operations fixes the mechanics in your head in a way no high-level API does. The core is small, mathematical, and entirely made of matrix multiplication, which is exactly why GPUs matter so much for it. The part worth carrying forward is that a token's representation is not a fixed property of the token. It is recomputed as a blend of every other token, weighted by learned relevance, on every forward pass. That is the step up from fixed-context models. - **Write the shapes down before the code.** Every bug in this post shows up as a shape you did not expect, or a plausible shape computing the wrong thing. - **Prefer the vectorized form over the loop.** The mechanism is matrix multiplication end to end, and that is what makes it parallel. - **Treat attention as a selector, not a language feature.** Selective information processing is the general capability; deciding which features matter to a specialization is what an expert in an MoE has to do. Next, we extend this single-head attention into **Multi-Head Attention** for more parallelism and richer context, the last piece before the full Transformer block. > Context is not stored in a token. It is a weighted sum, recomputed for every token, every time. --- Source: https://himanshuat.com/blogs/understanding-transformers-part-1-attention-from-scratch ======================================================================== --- title: "Part 2: Assembling the Full Architecture - From Attention to a Working Model" description: "Attention on its own is a lookup. A model is what you get when you wire lookups into a residual stream — positional encoding, masking, cross-attention — and check the tensor shape at every hop." date: "May 17, 2024" url: "https://himanshuat.com/blogs/understanding-transformers-part-2-full-architecture" --- # Part 2: Assembling the Full Architecture - From Attention to a Working Model Wire the decoder's self-attention exactly the way you wired the encoder's. Everything still runs. Same shapes, same code path, no error anywhere. The decoder has just been allowed to attend to the tokens it is supposed to predict, and you find out at generation time, when there is no future left to read. Nothing in the tensor shapes tells you the mask is missing. The shapes are identical either way. > Attention is the part everyone explains. What turns it into a model is the wiring around it — the mask, the residual, the norm — and the wiring is where a working architecture goes wrong. Part 1 worked through Multi-Head Attention, the mechanism that let us look at whole sequences without the RNN bottleneck. Attention on its own is a sophisticated lookup. To build something that processes information the way a brain does, which is my longer-term goal with MoE, you have to wire those lookups into a larger system. For sequence modeling that system is the Transformer. No skimming, and no LangChain abstraction hiding the computation. Component by component: the mechanics, the data flow, and why each piece is there. Getting the tensor shapes right is most of the work. ### Chapter 0 — What a block actually is A Transformer block is not a pipeline. It is a **residual stream** with sub-layers writing into it. The stream starts as the embedding. Each sub-layer reads the stream, computes something, and adds its result back: `sublayer_output + input`. The stream itself is never replaced, only added to. That single detail is what lets gradients flow directly through the network and makes depth survivable. An encoder block has two sub-layers, attention and a feed-forward network. A decoder block has three, because it adds a second attention that looks at the encoder instead of at itself. Everything else in this post is those two shapes, repeated `N` times. For me AI is less about building cool tech and more about understanding intelligence well enough to replicate it. The brain, with its specialized regions and its parallelism, is the reference. RNNs modelled this sequentially, forcing information through a bottleneck and struggling with long-range dependencies. The Transformer changed the shape of the problem: context processed in parallel. That is a step toward the modular, expert-driven architectures I want to build, and to use it well you need to know every bolt and wire. --- ### Step 1 — Give the sequence back its order Raw tokens map to dense vectors first. That part is standard. The addition that matters is **Positional Encoding**. Attention processes all tokens in parallel, so it has no idea which token came first. PE injects that information back. The sinusoidal form lets the model learn relative positions and generalize to sequences longer than the ones it trained on. ```typescript // Assuming 'embeddings' is your tensor of shape (batch_size, seq_len, d_model) // Let's use pseudo-TypeScript for clarity, implying raw tensor operations. class PositionalEncoding { private positionTable: Float32Array; // Precomputed table constructor(d_model: number, max_seq_len: number = 2048) { // Initialize sin/cos table based on d_model and max_seq_len // (Simplified for brevity, actual calculation involves 1/10000^(2i/d_model)) this.positionTable = new Float32Array(max_seq_len * d_model); // ... Populate positionTable with sinusoidal values ... } apply(embeddings: Tensor): Tensor { // Tensor: (batch_size, seq_len, d_model) const [batchSize, seqLen, d_model] = embeddings.shape; if (seqLen > this.positionTable.length / d_model) { console.warn("Sequence length exceeds max_seq_len for PE. Dynamic resizing or error handling needed."); } // Slice relevant part of precomputed PE table const peSlice = new Tensor( this.positionTable.slice(0, seqLen * d_model), // Get (seq_len, d_model) [seqLen, d_model] ); // Add PE to embeddings. Broadcasting handles batch_size. // Tensor operation: embeddings + peSlice return embeddings.add(peSlice); } } ``` The table is precomputed once and sliced per call, so the cost per forward pass is one broadcast add. The shapes: | tensor | shape | |---|---| | Input tokens | `(batch_size, seq_len)` | | Input embeddings | `(batch_size, seq_len, d_model)` | | Positional encoding | `(1, seq_len, d_model)` — broadcasts across batch | | Output | `(batch_size, seq_len, d_model)` | **The tell:** if you can shuffle the input tokens and the output is unchanged, positional encoding is not being applied. --- ### Step 2 — Wire attention, then a per-position FFN Multi-Head Attention is the core of the block. In the encoder, Q, K and V all come from the same source, which is what makes it self-attention. ```typescript // Assuming a MultiHeadAttention class from Part 1. // Inputs: query, key, value - all of shape (batch_size, seq_len, d_model) // (In the Encoder, Q, K, V all come from the same source) // Outputs: (batch_size, seq_len, d_model) class MultiHeadAttention { // ... constructor and selfAttention method from Part 1 ... // selfAttention(q: Tensor, k: Tensor, v: Tensor, mask: Tensor | null = null): Tensor { ... } } ``` Attention mixes information *across* positions. The feed-forward network is the opposite: it processes each position independently, with no reference to any other token in the sequence. Two linear layers with a non-linearity between them, widening to `d_ff` (usually `4 * d_model`) and coming back down. ```typescript class FeedForward { private layer1: LinearLayer; // d_model -> d_ff (usually 4 * d_model) private layer2: LinearLayer; // d_ff -> d_model constructor(d_model: number, d_ff: number, activation: (x: Tensor) => Tensor = relu) { this.layer1 = new LinearLayer(d_model, d_ff); this.layer2 = new LinearLayer(d_ff, d_model); this.activation = activation; } forward(input: Tensor): Tensor { // input: (batch_size, seq_len, d_model) // layer1: (batch_size, seq_len, d_model) -> (batch_size, seq_len, d_ff) let output = this.activation(this.layer1.forward(input)); // layer2: (batch_size, seq_len, d_ff) -> (batch_size, seq_len, d_model) output = this.layer2.forward(output); return output; } } ``` The widening to `d_ff` and back is the only place in the block where the last dimension changes: | stage | shape | |---|---| | Attention input (Q, K, V) | `(batch_size, seq_len, d_model)` | | Attention output | `(batch_size, seq_len, d_model)` | | FFN layer 1, pre-activation | `(batch_size, seq_len, d_ff)` | | FFN layer 2 output | `(batch_size, seq_len, d_model)` | **The tell:** if your FFN's forward pass touches the `seq_len` axis at all, you have written a second attention layer by accident. It should only ever see the last dimension. --- ### Step 3 — Wrap every sub-layer in a residual and a norm These two components rarely get top billing. Without them, training a network this deep runs straight into vanishing and exploding gradients. **Residual connections** add each sub-layer's output back to its input, which gives gradients an uninterrupted path from the loss to the earliest layers. **Layer normalization** normalizes activations across the feature dimension, for each example independently. Unlike BatchNorm, it works per example and per time step rather than across the batch, which is what makes it usable with variable sequence lengths and mixed-length batches. ```typescript // Let's assume a Tensor library with `add` and `mean`, `variance`, `div`, `mul`, `sub` operations. // LayerNorm applies across the last dimension (d_model). class LayerNormalization { private gamma: Tensor; // trainable scale parameter (d_model) private beta: Tensor; // trainable shift parameter (d_model) private epsilon: number = 1e-6; // small value for numerical stability constructor(d_model: number) { this.gamma = Tensor.ones([d_model]); // Initialize with ones this.beta = Tensor.zeros([d_model]); // Initialize with zeros } forward(input: Tensor): Tensor { // input: (batch_size, seq_len, d_model) const [batchSize, seqLen, d_model] = input.shape; // Calculate mean and variance across the last dimension (d_model) // Keep dimensions for broadcasting const mean = input.mean(2, true); // (batch_size, seq_len, 1) const variance = input.variance(2, true); // (batch_size, seq_len, 1) // Normalize: (input - mean) / sqrt(variance + epsilon) const normalized = input.sub(mean).div(variance.add(this.epsilon).sqrt()); // Apply learnable gamma and beta: gamma * normalized + beta // gamma and beta are (d_model), will broadcast to (batch_size, seq_len, d_model) return normalized.mul(this.gamma).add(this.beta); } } // How Add&Norm looks in practice: // x = input; // attention_output = this.attention.forward(this.norm1.forward(x)); // x = x.add(attention_output); // Residual connection // ffn_output = this.ffn.forward(this.norm2.forward(x)); // x = x.add(ffn_output); // Residual connection ``` Note the ordering in that last comment block, because it is the thing most people get backwards when reading the diagram in the paper. The norm is applied to what goes *into* the sub-layer. What gets added back is the raw, un-normalized input. The residual path stays untouched from one end of the stack to the other. `epsilon` at `1e-6` is there so a position whose features happen to be near-constant does not divide by something close to zero. | tensor | shape | |---|---| | Input | `(batch_size, seq_len, d_model)` | | Mean / variance, `keep_dims=true` | `(batch_size, seq_len, 1)` | | Normalized output | `(batch_size, seq_len, d_model)` | **The tell:** if you can delete a sub-layer from your block and the output shape changes, you have a pipeline, not a residual stream. In a residual stream, every sub-layer is optional at the shape level and mandatory at the quality level. --- ### Step 4 — Assemble the encoder block Two sub-layers, each wrapped the same way. That is the whole encoder. ```typescript class EncoderBlock { private selfAttention: MultiHeadAttention; private feedForward: FeedForward; private norm1: LayerNormalization; private norm2: LayerNormalization; constructor(d_model: number, n_heads: number, d_ff: number) { this.selfAttention = new MultiHeadAttention(d_model, n_heads); this.feedForward = new FeedForward(d_model, d_ff); this.norm1 = new LayerNormalization(d_model); this.norm2 = new LayerNormalization(d_model); } forward(input: Tensor, mask: Tensor | null = null): Tensor { // input: (batch_size, seq_len, d_model) // Sub-layer 1: Multi-Head Self-Attention + Add&Norm let attentionInput = this.norm1.forward(input); // (batch_size, seq_len, d_model) let attentionOutput = this.selfAttention.selfAttention( attentionInput, attentionInput, attentionInput, mask ); // (batch_size, seq_len, d_model) let output = input.add(attentionOutput); // Residual: (batch_size, seq_len, d_model) // Sub-layer 2: Feed-Forward + Add&Norm let ffnInput = this.norm2.forward(output); // (batch_size, seq_len, d_model) let ffnOutput = this.feedForward.forward(ffnInput); // (batch_size, seq_len, d_model) output = output.add(ffnOutput); // Residual: (batch_size, seq_len, d_model) return output; } } ``` Drawn out, the fork-and-merge is the shape of the block: ```mermaid flowchart TD X[input] --> N1[LayerNorm] N1 --> A[self-attention] X --> R1((add)) A --> R1 R1 --> N2[LayerNorm] N2 --> F[feed-forward] R1 --> R2((add)) F --> R2 R2 --> O[output] ``` Input and output are both `(batch_size, seq_len, d_model)`, which is exactly why blocks stack. The block is shape-preserving by construction, so `N` of them compose without a single adapter layer between them. **The tell:** if your encoder block's output shape differs from its input shape, you cannot stack it, and you have built one layer rather than an architecture. --- ### Step 5 — Mask the decoder so it cannot read its own answer The decoder generates an output sequence one token at a time, from the encoded input and the tokens it has already produced. Its first sub-layer is self-attention, identical to the encoder's except for one thing: **masking**. A causal, look-ahead mask stops the decoder from attending to future tokens during training. The prediction for position `i` depends only on positions `0` through `i-1`. ```typescript // When calling selfAttention: // this.selfAttention.selfAttention(attentionInput, attentionInput, attentionInput, causalMask); // causalMask is a triangular matrix of negative infinities for future positions. ``` Negative infinity rather than zero, because the mask is applied before the softmax. A zero would still leave a finite score to normalize against. Negative infinity is what makes the resulting attention weight exactly zero. **The tell:** feed the decoder a sequence, then change only the last token and re-run. If the hidden state at position 0 moves, your mask is not being applied — nothing at position 0 is allowed to depend on anything later. --- ### Step 6 — Feed the encoder's output in as keys and values Cross-attention is where the decoder actually reads the input. The wiring is the whole idea: - **Query** comes from the decoder's masked self-attention output. - **Key** and **Value** come from the encoder's output. The decoder asks the questions, the encoder holds the answers. That is what lets each generated token focus on the relevant part of the source sequence. ```typescript // Assuming `encoderOutput` is from the final Encoder layer. // And `decoderSelfAttentionOutput` is from the masked MHA layer in the decoder. // this.crossAttention.selfAttention( // decoderSelfAttentionOutput, // Query // encoderOutput, // Key // encoderOutput // Value // ); ``` This is the one place in the architecture where two different sequence lengths meet in the same operation: | tensor | masked self-attention | cross-attention | |---|---|---| | Q | `(batch_size, target_seq_len, d_model)` | `(batch_size, target_seq_len, d_model)` | | K, V | `(batch_size, target_seq_len, d_model)` | `(batch_size, source_seq_len, d_model)` | | Output | `(batch_size, target_seq_len, d_model)` | `(batch_size, target_seq_len, d_model)` | The output length follows Q, not K. Three sub-layers, three norms: ```typescript class DecoderBlock { private maskedSelfAttention: MultiHeadAttention; private crossAttention: MultiHeadAttention; private feedForward: FeedForward; private norm1: LayerNormalization; private norm2: LayerNormalization; private norm3: LayerNormalization; constructor(d_model: number, n_heads: number, d_ff: number) { this.maskedSelfAttention = new MultiHeadAttention(d_model, n_heads); this.crossAttention = new MultiHeadAttention(d_model, n_heads); this.feedForward = new FeedForward(d_model, d_ff); this.norm1 = new LayerNormalization(d_model); this.norm2 = new LayerNormalization(d_model); this.norm3 = new LayerNormalization(d_model); } forward( targetInput: Tensor, // (batch_size, target_seq_len, d_model) encoderOutput: Tensor, // (batch_size, source_seq_len, d_model) targetMask: Tensor | null = null, sourceMask: Tensor | null = null // For padding in encoder output ): Tensor { // Sub-layer 1: Masked Multi-Head Self-Attention + Add&Norm let selfAttentionInput = this.norm1.forward(targetInput); let selfAttentionOutput = this.maskedSelfAttention.selfAttention( selfAttentionInput, selfAttentionInput, selfAttentionInput, targetMask ); let output = targetInput.add(selfAttentionOutput); // Sub-layer 2: Multi-Head Cross-Attention (Encoder-Decoder Attention) + Add&Norm let crossAttentionInput = this.norm2.forward(output); let crossAttentionOutput = this.crossAttention.selfAttention( crossAttentionInput, // Query from decoder encoderOutput, // Key from encoder encoderOutput, // Value from encoder sourceMask // Mask for encoder padding ); output = output.add(crossAttentionOutput); // Sub-layer 3: Feed-Forward + Add&Norm let ffnInput = this.norm3.forward(output); let ffnOutput = this.feedForward.forward(ffnInput); output = output.add(ffnOutput); return output; } } ``` Same residual stream as the encoder, with one more merge in it: ```mermaid flowchart TD T[target input] --> N1[LayerNorm] N1 --> S[masked self-attention] T --> R1((add)) S --> R1 R1 --> N2[LayerNorm] N2 --> C[cross-attention] E[encoder output] -->|K, V| C R1 --> R2((add)) C --> R2 R2 --> N3[LayerNorm] N3 --> F[feed-forward] R2 --> R3((add)) F --> R3 R3 --> O[output] ``` Two masks are in play and they do different jobs. `targetMask` is causal and hides the future. `sourceMask` hides padding in the encoder output. Passing one where the other belongs is a bug that trains without complaining. **The tell:** if swapping `encoderOutput` for the decoder's own state in the K and V slots still runs, your cross-attention has no dependency on the encoder and you have built two disconnected stacks. --- ### Step 7 — Stack N of each and project to the vocabulary Then stack `N` encoder blocks and `N` decoder blocks, and put a linear layer on the end to get from `d_model` to vocabulary size. ```typescript class Transformer { private encoder: Encoder; // Stack of EncoderBlocks private decoder: Decoder; // Stack of DecoderBlocks private targetOutputLayer: LinearLayer; // To project to vocabulary size constructor( vocab_size: number, d_model: number, n_heads: number, d_ff: number, n_encoder_layers: number, n_decoder_layers: number, max_seq_len: number ) { this.encoder = new Encoder(d_model, n_heads, d_ff, n_encoder_layers, max_seq_len); this.decoder = new Decoder(d_model, n_heads, d_ff, n_decoder_layers, max_seq_len); this.targetOutputLayer = new LinearLayer(d_model, vocab_size); } forward( sourceInput: Tensor, // (batch_size, source_seq_len) - raw tokens targetInput: Tensor, // (batch_size, target_seq_len) - raw tokens sourceMask: Tensor | null = null, // for padding in source targetMask: Tensor | null = null // causal mask for target, and padding ): Tensor { // Embeddings and Positional Encodings are handled within Encoder/Decoder for brevity // (Typically, a shared embedding layer is used, with separate PEs) // Encoder pass const encoderOutput = this.encoder.forward(sourceInput, sourceMask); // (batch_size, source_seq_len, d_model) // Decoder pass let decoderOutput = this.decoder.forward( targetInput, encoderOutput, targetMask, sourceMask ); // (batch_size, target_seq_len, d_model) // Project decoder output to vocabulary size for token prediction // (batch_size, target_seq_len, d_model) -> (batch_size, target_seq_len, vocab_size) const finalLogits = this.targetOutputLayer.forward(decoderOutput); return finalLogits; // Raw logits before softmax } } ``` The final shape is `(batch_size, target_seq_len, vocab_size)`, and those are raw logits. Softmax belongs to the loss function or the sampler, not to the model. **The tell:** if your model's forward pass returns probabilities, you will end up applying softmax twice the first time you plug it into a cross-entropy loss. --- ### What breaks it **The mask that isn't there.** Nothing about a missing causal mask changes a tensor shape, so no runtime error fires. Training looks fine, because the decoder can read the answer. Generation is where it shows up, and by then you are debugging the sampler. **The sequence outruns the precomputed table.** The PE table is built once for `max_seq_len`, defaulting to 2048. Feed a longer sequence and the code above emits a `console.warn` and then slices anyway, which is exactly the kind of warning nobody sees in a training log. A hard failure here is friendlier than a quiet one. **Depth without the glue.** Residual connections and layer normalization are what make a deep stack trainable at all. Take them out and gradients vanish or explode long before any of the attention machinery gets a chance to matter. The components get the credit; the scaffolding does the work. **Normalizing across the batch.** Reaching for BatchNorm out of habit couples every example to the others in its batch. With variable sequence lengths, that statistic is being computed over padding as much as over content. LayerNorm's per-example, per-time-step scope is the reason it is the one used here. --- ### When the full encoder-decoder Transformer is the wrong choice Building the whole thing is the right way to understand it. It is often the wrong thing to deploy. - **You only need a representation.** Classification, embedding, retrieval. The encoder stack alone produces the contextual representation you want, and the decoder plus its cross-attention are parameters you train, store, and never read from. - **There is no separate source sequence.** If you generate from a prompt rather than translating one sequence into another, cross-attention has nothing distinct to attend to — K and V would come from the same stream as Q. A decoder-only stack is the honest shape for that job. - **You are shipping, not learning.** Working at the raw tensor level gives direct control, and it also means you own the performance work. This is the layer where production reaches for optimized CUDA kernels or WebGPU compute shaders to make the matmuls and element-wise ops fast. The framework abstraction that I deliberately refused here exists precisely because someone already did that. **The tell is Step 6.** If you cannot name what goes in the K and V slots that isn't already in Q, you do not need an encoder-decoder model. --- ### The shift The Transformer's real lesson is not attention. It is that a few simple components — attention, a feed-forward network — connected by residuals and normalization, stack into something capable without any adapter glue between the layers. Shape-preserving blocks compose. That property is the architecture. That modularity is also exactly what MoE needs. Build one expert from a Transformer block, then orchestrate many of them, and the interesting problem moves from what a block computes to which blocks run at all. - **Trace `(batch_size, seq_len, d_model)` through every single operation.** One mismatch costs an hour, and the trace is also how you actually learn the data flow. - **Treat the block as a residual stream, not a pipeline.** Sub-layers add to it; they do not replace it. - **Assume the mask is missing until you have proven it isn't.** It is the one bug in this architecture that shapes will never catch for you. What I want to figure out next is how to assemble these blocks into a system that allocates compute dynamically, the way a brain activates specific regions for a task. Replicating even a fraction of that is a long road, and a Transformer understood bolt by bolt is the right place to start it. --- Source: https://himanshuat.com/blogs/understanding-transformers-part-2-full-architecture ======================================================================== --- title: "Part 3: The Scaling Problem - Optimizing Transformer Memory and Compute" description: "At N=8192 the attention score matrix is 67 million elements, about 268 MB per head per layer in float32. The math is fine. The memory traffic is what kills you, and FlashAttention, quantization, and LoRA each attack a different part of that bill." date: "May 24, 2024" url: "https://himanshuat.com/blogs/understanding-transformers-part-3-optimization-and-scaling" --- # Part 3: The Scaling Problem - Optimizing Transformer Memory and Compute Set the context window to `N=8192` and the attention score matrix `Q @ K^T` becomes 67 million elements. In `float32` that is roughly 268 MB — per head, per layer. Multiply by heads and layers and the GPU's HBM saturates before you have done anything interesting with the model. Vanilla self-attention is mathematically clean. Clean math and efficient computation are different problems, and the second one is what stops you. > The quadratic term in attention is a memory-traffic bill before it is a FLOPs bill. You do not fix it by rewriting the formula. You fix it by changing where the intermediate values live. ### Chapter 0 — The two bottlenecks, which are not the same bottleneck A standard Transformer hits the wall in two separate places, and conflating them is why people reach for the wrong fix. **Memory complexity, `O(N^2)` in activations.** Storing the score matrix and the softmax output over it costs `O(N^2)` memory in the sequence length `N`. That is the 268 MB per head per layer above. The KV cache used at inference scales quadratically too. **Compute complexity, `O(N^2)` in the attention matmuls.** `Q @ K^T` and `softmax(scores) @ V` both scale quadratically. GPUs are fast at matrix ops, but `N^2` catches up, and even with memory to spare the raw FLOPs become prohibitive. These are what cap the context window, and long context is exactly what multi-turn dialogue and long-form understanding need. Making a big model run at all is the low bar. The bar that matters is letting it work over more context. --- ### Step 1 — Move attention into SRAM instead of HBM The `O(N^2)` problem is partly about the number of operations and mostly about *where* those operations touch the memory hierarchy. Compute cores are very fast. Moving data between slow High Bandwidth Memory and the small, fast on-chip SRAM is the real bottleneck. FlashAttention, from Tri Dao et al., does not change the mathematical output of attention. It reorders the computation to minimise HBM reads and writes. The naive path computes the full `Q @ K^T`, applies `softmax`, and multiplies by `V` — which forces the intermediate `N^2` matrices out to HBM and back. FlashAttention tiles instead. It splits `Q` and `K` into blocks, computes attention block by block, and accumulates the `softmax(QK^T)V` output in SRAM without ever writing the full `QK^T` matrix out. The softmax reduction runs iteratively over blocks, which is what keeps it numerically stable while streaming. ```mermaid flowchart LR A[next Q,K,V block] --> B[scores in SRAM] B --> C[rescale m_i, l_i] C --> D[accumulate output] D -->|blocks remain| A D -->|done| E[write N x D to HBM] ``` The running max `m_i` and normaliser `l_i` are the whole trick: they let a partial softmax computed on one block be corrected once a later block reveals a larger max. ```python # Conceptual Pseudo-code for FlashAttention's core idea (not runnable CUDA) # Imagine this operating directly on GPU registers and shared memory. def flash_attention_blockwise(Q_block, K_block, V_block, prev_l_i, prev_m_i, output_accumulator): """ Simulated FlashAttention-like block processing. In reality, this is a highly optimized CUDA kernel. """ # 1. Load Q_block, K_block, V_block into fast SRAM. # (Simulated by direct access here for conceptual clarity) # 2. Compute attention scores for this block in SRAM S_block = Q_block @ K_block.T # O(block_size^2) local ops # 3. Apply part of softmax calculation iteratively # This involves managing normalizer (l_i) and max (m_i) in a streaming fashion. m_i_new = max(prev_m_i, S_block.max()) P_block = exp(S_block - m_i_new) # Scale for numerical stability l_i_new = prev_l_i * exp(prev_m_i - m_i_new) + P_block.sum() # Update normalizer # 4. Compute partial output O_block = P_block @ V_block # 5. Combine with previous accumulated output, weighted by normalizers output_accumulator = (prev_l_i * exp(prev_m_i - m_i_new) * output_accumulator + O_block) / l_i_new return output_accumulator, l_i_new, m_i_new # In practice, you'd just use a library call like this: # from flash_attn import flash_attn_func # output = flash_attn_func(q, k, v, dropout_p=0.0, causal=True) ``` You would not hand-write this. You call `flash_attn_func`. The pseudo-code is here because the rescaling step is the part that makes the tiling legal, and it is worth being able to point at. Keeping the `N^2` intermediates in SRAM and writing only the final `N x D` output to HBM gives a typical 2-4x speedup and cuts HBM usage sharply. That is what makes `N=8K` or `N=16K` feasible where it was not. **The tell:** if your attention path materialises a full `N x N` score matrix as a tensor you could print, you are paying the HBM bill this kernel exists to delete. --- ### Step 2 — Spend fewer bits per weight A 7B parameter model in `float32` is roughly 28 GB of memory for the weights alone. Add activations, gradients, and the KV cache and you exceed what even high-end GPUs offer. Quantization reduces the precision of weights and activations from high-precision floats to lower-precision integers. It works because many weights do not need 32 bits of precision, and the saving compounds: less data to move can also speed up computation, since smaller matrices fit better in cache. Two approaches, and the choice is a straightforward trade: | approach | when quantization happens | accuracy | complexity | |---|---|---|---| | Post-Training Quantization (PTQ) | after training | can drop | simplest | | Quantization-Aware Training (QAT) | simulated during training | more accurate | more complex | For large models in production, especially inference, PTQ to INT8 or INT4 is common. `bitsandbytes` on PyTorch, or native framework support, handles the mechanics. ```python import torch def quantize_linear_layer(linear_layer, num_bits=8): """ Conceptual INT8 quantization for a single linear layer. In reality, this involves specific scaling factors per tensor/per group. """ if num_bits not in [4, 8]: raise ValueError("Only 4 or 8 bits supported for this demo.") # Get original FP32 weights weights_fp32 = linear_layer.weight.data # Calculate scale and zero_point (simplified for concept) # A common approach is symmetric quantization: s = max_abs_val / (2^(num_bits-1) - 1) # Or asymmetric: s = (max - min) / (2^num_bits - 1), z = -round(min/s) max_val = weights_fp32.abs().max() scale = max_val / (2**(num_bits - 1) - 1) # Symmetric quantization # Quantize: weights_int = round(weights_fp32 / scale) # Clamp to the target integer range (e.g., [-127, 127] for INT8) int_min = -(2**(num_bits - 1)) int_max = (2**(num_bits - 1)) - 1 weights_int = torch.clamp(torch.round(weights_fp32 / scale), int_min, int_max).to(torch.int8 if num_bits==8 else torch.int4) # Store quantized weights and quantization parameters linear_layer.quantized_weight = weights_int linear_layer.quant_scale = scale linear_layer.quant_zero_point = 0 # Symmetric print(f"Layer {linear_layer} quantized to INT{num_bits}.") print(f"Original size: {weights_fp32.nelement() * 4 / (1024**2):.2f} MB") print(f"Quantized size: {weights_int.nelement() * num_bits / 8 / (1024**2):.2f} MB") # For actual inference, you'd then de-quantize on the fly # or use specialized kernels for INT8/INT4 matrix multiplication. # E.g., output = (input_int @ weights_int) * scale_input * scale_weight / scale_output # Example Usage: # model = MyTransformerModel() # for name, module in model.named_modules(): # if isinstance(module, torch.nn.Linear): # quantize_linear_layer(module, num_bits=8) ``` The single `scale` derived from `weights_fp32.abs().max()` is the simplification to notice. One outlier weight in a tensor stretches the scale and costs every other weight resolution, which is why real implementations quantize per tensor or per group rather than per layer. Against FP32, the size reduction is fixed by the format: | precision | size vs FP32 | |---|---| | FP16 | 2x smaller | | INT8 | 4x smaller | | INT4 | 8x smaller | That is what lets you run larger models on consumer GPUs, or cut inference cost on specialised hardware. The cost is a usually small accuracy drop, paid up front and permanently — you are not renting the precision back later. **The tell:** if the model weights load and the process dies on the first forward pass, precision is the knob, and the weights were never the whole bill. --- ### Step 3 — Train a low-rank delta and freeze the base Fine-tuning something like LLaMA-2 70B for one task means updating tens of billions of parameters. That is expensive in compute and in memory, since gradients and optimizer states scale with the parameters you are updating. It also leaves you storing a full copy of the model per task, which does not scale to the number of tasks a general-purpose system needs. PEFT methods update a small subset of parameters, or introduce a few new trainable ones, while the vast majority of the pre-trained model stays frozen. **LoRA (Low-Rank Adaptation)**, from Hu et al., 2021, is the clean version of this. For each pre-trained weight matrix `W_0`, it approximates the update `ΔW` as a low-rank decomposition `ΔW = BA`: - `W_0` — original pre-trained weight matrix, `d_out x d_in`. Frozen. - `A` — `r x d_in`, trainable. - `B` — `d_out x r`, trainable. - `r` — the rank, typically 4, 8, or 16. The forward pass becomes `h = W_0 x + BA x`. Only `A` and `B` are trained. ```python import torch import torch.nn as nn class LoRALinear(nn.Module): def __init__(self, original_linear_layer, rank=8, alpha=16): super().__init__() self.original_linear = original_linear_layer self.original_linear.weight.requires_grad = False # Freeze original weights self.original_linear.bias.requires_grad = False # Freeze original bias in_features = original_linear_layer.in_features out_features = original_linear_layer.out_features # LoRA A and B matrices self.lora_A = nn.Parameter(torch.randn(rank, in_features)) self.lora_B = nn.Parameter(torch.randn(out_features, rank)) # Scaling factor, a common practice in LoRA self.scaling = alpha / rank # Initialize LoRA B to zeros for identity initialization (no change initially) nn.init.zeros_(self.lora_B) def forward(self, x): # Original forward pass with frozen weights original_output = self.original_linear(x) # LoRA adaptation lora_output = (self.lora_B @ self.lora_A) @ x.T lora_output = lora_output.T * self.scaling # Apply scaling and transpose back return original_output + lora_output # Example Usage: # original_linear = nn.Linear(768, 768) # e.g., in a self-attention projection # lora_layer = LoRALinear(original_linear, rank=8, alpha=16) # # Now, only lora_layer.lora_A and lora_layer.lora_B are trainable parameters # for name, param in lora_layer.named_parameters(): # print(f"{name}: requires_grad={param.requires_grad}") # # Output: # # lora_A: requires_grad=True # # lora_B: requires_grad=True ``` `nn.init.zeros_(self.lora_B)` is the line that makes this safe to attach to a trained model. With `B` at zero, `BA` is zero, and the adapted layer is bit-for-bit the original at step 0. Training starts from the base model's behaviour rather than from noise injected into it. Those two small matrices cut trainable parameters sharply — often more than 1000x versus full fine-tuning. Less memory, faster training, and you keep many task-specific adapters (just `A` and `B`) around one shared base. For a mixture-of-experts design where many experts share a base and adapt to context, that shape is exactly the one you want. **The tell:** if you are storing a full model checkpoint per task, you are paying to store parameters that never changed between them. --- ### What breaks it **FlashAttention removes the memory bottleneck and leaves the compute one standing.** The two bottlenecks in Chapter 0 are independent. Tiling into SRAM deletes the `N^2` HBM traffic; it does not delete the `N^2` matmuls. Push `N` far enough and the FLOPs become prohibitive on their own, with plenty of memory free. Anyone who reports the 2-4x speedup as though it made attention linear has read the result wrong. **Quantization's accuracy cost is unrecoverable at inference.** PTQ is the simple path and it is the one that can drop accuracy, because the model never saw the reduced precision during training. QAT buys the accuracy back by simulating quantization in the training loop, which means you need the training loop — not an option if you are quantizing someone else's released weights. **A rank-8 delta can only express a rank-8 update.** `ΔW = BA` with `r` at 4, 8, or 16 is a strong structural constraint on what fine-tuning is allowed to change. That is the source of the >1000x parameter saving and it is also the limit: an update that genuinely is not low-rank cannot be represented, no matter how long you train `A` and `B`. --- ### When these optimizations are the wrong choice Each of these buys headroom at a price, and the price is only worth paying under specific conditions. - **Your sequences are short.** The whole argument rests on `N^2` dominating. At small `N`, the quadratic term is not what is costing you, and a fused attention kernel is complexity with no payoff attached. - **You are FLOP-bound, not memory-bound.** If HBM is not saturated, reordering memory traffic optimises something that was never the constraint. Measure which wall you are actually at before picking the tool. - **Accuracy is the binding constraint.** Quantization trades precision for memory. If the model is already marginal on the task, INT4 spends the wrong currency, and PTQ in particular offers no way to earn it back. - **You are training one model for one task, forever.** PEFT's payoff is many adapters around one frozen base. With a single task and the memory to fine-tune fully, LoRA's rank constraint is a restriction you accepted for a benefit you never collect. --- ### The shift The architecture work reads well on paper: MoE, dynamic routing, symbolic reasoning. What determines whether any of it runs is the memory hierarchy underneath it. Performance here is not polish applied after the design. FlashAttention did not only make attention faster; it made long context windows feasible at all. Quantization is what puts a large base model on hardware you own. PEFT is what makes specialisation cheap enough to do many times. - **Find out which wall you are at.** Memory traffic and FLOPs are separate bottlenecks with separate fixes. - **Change where the intermediates live before you change the math.** Tiling into SRAM is an exact reordering. It costs no accuracy. - **Freeze what you are not changing.** A low-rank delta plus a shared base beats a checkpoint per task. Every byte saved and cycle gained is context the model gets to reason over instead. --- Source: https://himanshuat.com/blogs/understanding-transformers-part-3-optimization-and-scaling ======================================================================== --- title: "Part 4: An Architectural Deep Dive - Why BERT and GPT Are Different Beasts" description: "BERT and GPT run on the same self-attention machinery. One triangular matrix of negative infinities decides whether a model reads or writes, and every other difference between them follows from it." date: "May 31, 2024" url: "https://himanshuat.com/blogs/understanding-transformers-part-4-architecture-bert-vs-gpt" --- # Part 4: An Architectural Deep Dive - Why BERT and GPT Are Different Beasts GPT's attention calculation has one step BERT's doesn't. Every score above the diagonal gets set to negative infinity before the softmax. That single step decides whether the model can read a sentence or write one. Treat the two as interchangeable black boxes, especially through a framework that hides the internals, and you find out the moment you need to optimize for a specific task and have no idea which knob you are turning. > BERT and GPT run on the same underlying machinery. The only thing that differs is *which* other tokens a given token is allowed to attend to, and the pre-training objective, the task fit, and the failure modes all follow from that one choice. Bigger models aren't the whole story. Architecture decides how a model handles information, and those choices matter as much as scale. ### Chapter 0 — What the mask actually does The Transformer processes a sequence by letting each token weigh the importance of the other tokens, through self-attention over Query (Q), Key (K), and Value (V) matrices. A single attention head is five operations. Compute scores as `QK^T`. Scale them by `1 / sqrt(head_dim)`. Apply a mask. Softmax the result into weights. Multiply the weights by V. The mask sits at step three, and it is the only step where BERT and GPT disagree. An entry set to `-Infinity` before the softmax is blocked, because the softmax that follows gives it no weight. **BERT masks padding. GPT masks padding and the future.** Everything else in this post is a consequence of that line. --- ### Step 1 — Give the encoder the whole sentence BERT (Bidirectional Encoder Representations from Transformers) is an encoder-only model: a stack of Transformer encoder blocks, with no decoder and no cross-attention to a separate encoder. Each block takes a sequence of tokens and processes them with a full view of the entire input. Its attention is bidirectional. A token at position `t` can attend to every token from `1` to `N` (the sequence length): tokens to its left, itself, and tokens to its right. That two-way view is what lets it read context deeply, jumping between any words in a sentence to work out how they relate. Here is how to think about a single self-attention head inside a BERT-like encoder layer, written out in raw TypeScript. Note the absence of a causal mask; the only masking needed is for padding tokens. ```typescript // Assuming 'Matrix' is a simple 2D array or custom matrix class // and 'matmul', 'transpose', 'scale', 'softmax' are available low-level ops. interface Matrix { data: number[][]; dims: [number, number]; // [rows, cols] // ... basic matrix operations } // Low-level operation to apply a mask (e.g., for padding) function applyMask(scores: Matrix, mask: Matrix, invalidValue: number = -Infinity): Matrix { const maskedScores = JSON.parse(JSON.stringify(scores)); // Deep copy for (let i = 0; i < scores.dims[0]; i++) { for (let j = 0; j < scores.dims[1]; j++) { if (mask.data[i][j] === 0) { // Assuming 0 in mask means 'ignore' or 'pad' maskedScores.data[i][j] = invalidValue; } } } return maskedScores; } /** * Calculates bidirectional self-attention scores for a single head. * In a BERT encoder, tokens can attend to all other tokens. * @param query Query matrix (seq_len, head_dim) * @param key Key matrix (seq_len, head_dim) * @param value Value matrix (seq_len, head_dim) * @param paddingMask Optional mask for padding tokens (seq_len, seq_len), 0 where padded. * @returns Attention output matrix (seq_len, head_dim) */ function calculateBidirectionalAttention(query: Matrix, key: Matrix, value: Matrix, paddingMask?: Matrix): Matrix { // 1. Compute attention scores: QK^T const scores = matmul(query, transpose(key)); // Result: (seq_len, seq_len) // 2. Scale scores const headDim = key.dims[1]; const scaledScores = scale(scores, 1 / Math.sqrt(headDim)); // 3. Apply padding mask if provided let maskedScores = scaledScores; if (paddingMask) { maskedScores = applyMask(scaledScores, paddingMask, -Infinity); } // 4. Apply softmax to get attention weights const attentionWeights = softmax(maskedScores); // Each row sums to 1 // 5. Multiply weights by Value matrix return matmul(attentionWeights, value); // Result: (seq_len, head_dim) } ``` The `paddingMask` is optional in that signature for a reason. With no mask at all, the code still runs and still produces a valid encoder head. There is nothing structural stopping a token from looking right. The training objective is built on that freedom. BERT was trained mainly with masked language modeling (MLM): randomly mask out a percentage of tokens in a sentence, then train the model to predict the originals. Filling in a blank draws on context from both sides, so the objective requires the bidirectional view. It also used next sentence prediction (NSP) to learn how sentences relate. That combination puts BERT at its best on tasks that need deep comprehension of existing text: * Text classification: sentiment analysis, spam detection. * Named entity recognition (NER): identifying names, organizations, locations. * Extractive question answering: finding the exact answer span within a text. * Search and information retrieval: ranking relevant documents. * Extractive summarization: selecting the most important sentences. **The tell:** if the only mask your attention code builds is derived from token padding, you are running an encoder, and its output is a representation rather than a next token. --- ### Step 2 — Block the future with a triangular mask GPT (Generative Pre-trained Transformer) models are decoder-only: stacked Transformer decoder blocks that omit the encoder-decoder cross-attention layer from the original Transformer's decoder. Its attention is causal. A token at position `t` can attend only to tokens `1` through `t`; it cannot see anything that comes after it. A causal mask enforces this during the attention calculation. The one-way flow is what makes generation work, since you produce language one token at a time without access to the future. The mask itself is a triangular matrix: values above the diagonal are set to negative infinity, which blocks future tokens. ```typescript // Assuming previous Matrix interface and helper functions /** * Generates a causal mask matrix. * For a token at index `i`, it can only see tokens at index `j` where `j <= i`. * @param seqLen The sequence length. * @returns A (seqLen, seqLen) matrix where future positions are -Infinity. */ function generateCausalMask(seqLen: number): Matrix { const maskData: number[][] = Array(seqLen).fill(0).map(() => Array(seqLen).fill(0)); for (let i = 0; i < seqLen; i++) { for (let j = i + 1; j < seqLen; j++) { maskData[i][j] = -Infinity; // Block future tokens } } return { data: maskData, dims: [seqLen, seqLen] }; } /** * Calculates causal self-attention scores for a single head. * In a GPT decoder, tokens can only attend to previous tokens. * @param query Query matrix (seq_len, head_dim) * @param key Key matrix (seq_len, head_dim) * @param value Value matrix (seq_len, head_dim) * @param paddingMask Optional mask for padding tokens (seq_len, seq_len), 0 where padded. * @returns Attention output matrix (seq_len, head_dim) */ function calculateCausalAttention(query: Matrix, key: Matrix, value: Matrix, paddingMask?: Matrix): Matrix { // 1. Compute attention scores: QK^T const scores = matmul(query, transpose(key)); // Result: (seq_len, seq_len) // 2. Scale scores const headDim = key.dims[1]; const scaledScores = scale(scores, 1 / Math.sqrt(headDim)); // 3. Generate and apply the causal mask const causalMask = generateCausalMask(scores.dims[0]); let maskedScores = applyMask(scaledScores, causalMask, -Infinity); // 4. Apply padding mask if provided (combine with causal mask) if (paddingMask) { // This is a simplified merge. In reality, you'd apply -Infinity if *either* mask dictates it. // A proper `combineMasks` would involve element-wise max/min or logic. maskedScores = applyMask(maskedScores, paddingMask, -Infinity); } // 5. Apply softmax to get attention weights const attentionWeights = softmax(maskedScores); // 6. Multiply weights by Value matrix return matmul(attentionWeights, value); } ``` Line for line this is the encoder head with one insertion at step three, and the causal mask is not optional there. It is generated unconditionally from the sequence length, because the model's whole training setup depends on it. GPT is trained with next token prediction: given a sequence, predict the next token. The objective enforces the causal constraint on its own, since predicting the next word can only use words that have already appeared. That makes GPT a natural fit for generative tasks: * Text generation: articles, stories, code, poems. * Chatbots and conversational AI: coherent, contextually relevant responses. * Abstractive summarization: writing new summary text rather than extracting sentences. * Translation, when fine-tuned for it: producing target-language text from a source. * Code generation: syntactically correct, working snippets. **The tell:** if your attention code builds a mask from `seqLen` alone, with no reference to the input tokens, you are running a decoder and the future is already closed off. --- ### Step 3 — Choose from the direction of the task, not the model card The two architectures line up cleanly once you put them side by side. | | BERT | GPT | |---|---|---| | Blocks | Encoder-only | Decoder-only, no cross-attention | | Attention | Bidirectional: `1` to `N` | Causal: `1` to `t` | | Mask | Padding only | Padding plus triangular causal | | Pre-training | Masked language modeling, next sentence prediction | Next token prediction | | Output | An encoding of existing text | New text, one token at a time | | Suits | Classification, NER, extractive QA, retrieval, extractive summarization | Generation, chat, abstractive summarization, translation, code | So the decision runs off one question about your task, and the architecture falls out of the answer. ```mermaid flowchart TD A[Your task] --> B{Read or write?} B -->|Understand, encode| C[Bidirectional] B -->|Generate tokens| D[Causal] C --> E[Encoder-only, BERT] D --> F[Decoder-only, GPT] ``` If you need to *understand* and *encode* text, the two-way view is the reason to reach for BERT. If you need to *generate* new text, GPT's autoregressive approach is the one to reach for. **The tell:** write down whether your task consumes text or emits it. If you cannot answer that in one word, you are not yet in a position to pick an architecture, and a benchmark table will not rescue you. --- ### What breaks it **Two masks, merged carelessly.** The causal snippet above applies the padding mask on top of the already-masked scores, and the comment says exactly what is wrong with that: it is a simplified merge. A correct combination applies `-Infinity` if *either* mask dictates it, which means element-wise logic rather than a second pass. Mask-combining is where hand-rolled attention quietly goes wrong, because the output still has the right shape. **Keeping the objective and swapping the mask.** The mask and the pre-training objective are one decision, not two. Masked language modeling only makes sense when a token can see both sides of the blank, and next token prediction only makes sense when it cannot see ahead. Change one and the other stops meaning anything. **Names that describe what was removed.** "Decoder-only" does not mean the original Transformer's decoder. It means those blocks with the encoder-decoder cross-attention layer taken out. Read the label as a full decoder and you will go looking for an encoder that was never there. --- ### When the BERT-versus-GPT split is the wrong lens This distinction earns its keep when you are optimizing a specific task. Outside that, it is a vocabulary lesson. - **The high-level API is doing fine.** Leaning on a framework that hides all of this works until it doesn't. If your task is served and you're not tuning it, the internals are not the thing to spend an afternoon on. - **Both of your candidates are decoder-only.** The encoder/decoder axis tells you nothing about two models that sit on the same side of it. Whatever separates them is somewhere else. - **Your model is neither.** The original Transformer has an encoder, a decoder, and cross-attention between them. Mixture-of-experts models add specialization that this two-way split doesn't describe. A binary is not a taxonomy. - **You haven't defined the task yet.** The mask follows from whether you're consuming text or producing it. Without that answer, there is nothing for the architecture to follow from. --- ### The shift The differences between BERT and GPT run deeper than surface details. They are two distinct approaches, encoded directly into the attention mechanism and, following from that, into the pre-training objective. There's a loose parallel to the brain here, with different regions specializing in taking input in versus producing output. My longer-term interest is in replicating parts of how that works, particularly its specialized processing, which is part of what draws me to MoE architectures. A modular AI probably won't be one giant model so much as a set of specialized parts, each handling the piece of the job it's suited to. - **BERT reads.** Bidirectional attention takes in the full context of a passage, which suits interpretation and classification. - **GPT writes.** Causal attention produces text one token at a time, which suits open-ended generation. - **The mask is the decision.** The objective, the task fit, and the limits are downstream of it. Go find the line in your model's attention implementation where the mask gets applied, and read what it blocks. That line tells you more about what the model is for than the model card does. --- Source: https://himanshuat.com/blogs/understanding-transformers-part-4-architecture-bert-vs-gpt ======================================================================== --- title: "Part 5: The Architectural Frontier - Mamba, RAG, and the Future Beyond Attention" description: "Attention costs O(N^2), so the context window is capped by arithmetic rather than by anything the model can or cannot do. Mamba trades the full history for a fixed-size state, and RAG moves the knowledge out of the weights. Here is what each one buys and what it costs." date: "June 07, 2024" url: "https://himanshuat.com/blogs/understanding-transformers-part-5-future-architectures" --- # Part 5: The Architectural Frontier - Mamba, RAG, and the Future Beyond Attention Every token attends to every other token. Double the sequence and the attention cost quadruples. So the wall you hit is $O(N^2)$, not model capability. It caps how far the context window scales, and anything resembling memory needs very long sequences. Self-attention reset how we do sequence modeling, and for a while "is attention all you need?" had a clear answer. Push on context length, on real *online learning*, on *recurrence*, and the answer stops being clear. > Whatever comes after attention won't be a single math trick. It'll be architectures that scale, sitting next to systems that can pull in knowledge which keeps changing. I'm less interested in building the next big model than in understanding how intelligence works from the inside. The system I keep coming back to would mimic the brain's modularity, something like a large Mixture of Experts (MoE) where different parts handle different tasks and share what they learn. Both halves of this post are pieces of that: Mamba, which challenges the Transformer directly, and retrieval-augmented generation (RAG), a systems approach that grounds a model in outside knowledge. ### Chapter 0 — What a state is Mamba is built on state-space models (SSMs). SSMs aren't new. Control systems and signal processing have used them for decades. The idea is to model how a system evolves through a hidden *state*. For a continuous system: $$ \frac{dh}{dt} = Ah + Bx \\ y = Ch + Dx $$ Where $h$ is the hidden state, $x$ is the input, $y$ is the output, and $A, B, C, D$ are matrices defining the system. When discretized for sequential data, it becomes: $$ h_t = A h_{t-1} + B x_t \\ y_t = C h_t + D x_t $$ Read the second pair carefully, because everything below is a consequence of it. $h_t$ depends on $h_{t-1}$ and $x_t$. Nothing else. The history that attention keeps around has been compressed into one fixed-size vector, and the cost of a step no longer depends on how many tokens came before it. --- ### Step 1 — Trade the full history for a fixed-size state That recurrence runs in linear time, $O(N)$, because each step reads only the previous state and the current input. That is the efficiency gain, and the rest follows from it. | | Transformer | SSM / Mamba | |---|---|---| | cost in sequence length | $O(N^2)$ | $O(N)$ | | what one step reads | every previous token | previous state, current token | | streaming and real-time | attention struggles here | recurrent by construction | | long sequences | memory and inference cost climb | lower memory, faster inference | Linear scaling is the headline: no quadratic wall, so much longer context windows come within reach. **The tell:** if a step in your model reads anything besides the previous state and the current input, its cost is still a function of history length, whatever the block is called. --- ### Step 2 — Make the state update data-dependent A fixed $A$, $B$, $C$ gives you a recurrence that treats every token the same way. Mamba makes those matrices functions of the input instead, so the model can choose what to carry forward and what to forget based on context. For language that matters a lot, because most tokens are not worth keeping in the state. Selectivity is the expensive part, and Mamba's answer is a hardware-aware parallel scan that keeps it fast on modern accelerators. The real implementation lives in optimized CUDA kernels. Below is the recurrence underneath it, written out so the state update is visible. ```python import torch def conceptual_selective_scan( inputs: torch.Tensor, # Shape: (Batch, SeqLen, Dim) A_matrix: torch.Tensor, # Shape: (Dim, StateDim) -> conceptually simplified B_projections: torch.Tensor, # Shape: (SeqLen, Dim, StateDim) -> data-dependent B C_projections: torch.Tensor, # Shape: (SeqLen, StateDim, Dim) -> data-dependent C ) -> torch.Tensor: """ A highly simplified, conceptual representation of Mamba's selective scan. In reality, A, B, C are carefully constructed and data-dependent through linear layers and activation functions, then discretized and optimized. This demonstrates the recurrent state update. """ batch_size, seq_len, input_dim = inputs.shape state_dim = A_matrix.shape[1] # A_matrix is typically a fixed part, but B, C are dynamic # Initial hidden state for each item in batch hidden_state = torch.zeros(batch_size, state_dim, device=inputs.device) outputs = [] for t in range(seq_len): current_input = inputs[:, t, :] # (Batch, Dim) # In real Mamba, B and C are *derived* from current_input # via MLPs, making them data-dependent and allowing selectivity. # Here, we assume they are pre-computed / projected for each step. B_t = B_projections[t] # (Dim, StateDim) C_t = C_projections[t] # (StateDim, Dim) # dt is also data-dependent in Mamba, contributing to selectivity. # For simplicity, A is treated as fixed across steps here. A_t = A_matrix # (Dim, StateDim) - In real Mamba, this is more complex. # The conceptual "scan": iterate the sequence, updating state each step. # Real code batches these matmuls instead of looping over the batch. for b in range(batch_size): hidden_state[b] = ( torch.einsum("sd,d->s", A_t, hidden_state[b]) + # Simplified A*h torch.einsum("sd,d->s", B_t, current_input[b]) # Simplified B*x ) # Output generation: y_t = C*h_t outputs.append(torch.einsum("ds,s->d", C_t, hidden_state[b])) # Simplified C*h return torch.stack(outputs, dim=1) # (Batch, SeqLen, Dim) ``` *Note: the real Mamba uses optimized parallel-scan operations and a more careful parameterization of `A`, `B`, `C` to get both data-dependence and hardware efficiency. The code above is a simplification meant to show the recurrent state update.* Mamba isn't a curiosity. It's competitive with Transformers on long-context tasks, sometimes better, and noticeably faster. **The tell:** if `A`, `B`, and `C` are the same tensors at every timestep, you have a plain recurrence. Selectivity is the part where they come from `current_input`. --- ### Step 3 — Move the changing knowledge out of the weights Mamba is about *processing* information efficiently. It does not change what the model knows, and no model holds all knowledge or stays accurate without some live connection to the world. A plain LLM has two problems. Its knowledge stops at the training data, which is always somewhat stale, and without grounded facts it will confidently make things up. RAG gives the LLM real-time access to an external, up-to-date knowledge base. It sits alongside the sequence model rather than replacing it. Think of it as giving the model a library it can look things up in. > A model with a knowledge cut-off has no way to notice it has one. That is the part that costs you. **The tell:** ask the model about something that changed after its training data was collected. A confident, wrong answer is not an architecture problem, and no amount of scaling fixes it. --- ### Step 4 — Build the pipeline out of four calls, not a framework I'd rather have direct control here, so I'm not reaching for a framework like LangChain. A RAG pipeline is really just a few direct API calls with minimal dependencies. 1. Embed the query. Turn the user's question into a vector. 2. Retrieve. Use that vector to search a vector database for semantically similar documents. 3. Build the prompt. Add the retrieved documents to the LLM prompt as context. 4. Generate. The LLM answers *from that context*. Here's how you'd do it in TypeScript, using raw `fetch` for API calls: ```typescript // Define interfaces for clarity interface EmbeddingResponse { data: [{ embedding: number[]; index: number; object: string; }]; model: string; object: string; usage: { prompt_tokens: number; total_tokens: number; }; } interface ChatCompletionResponse { choices: [{ finish_reason: string; index: number; message: { content: string; role: string; }; }]; created: number; id: string; model: string; object: string; usage: { completion_tokens: number; prompt_tokens: number; total_tokens: number; }; } interface Document { id: string; content: string; // In a real system, you'd store the embedding as well // embedding: number[]; } // --- Configuration --- const OPENAI_API_KEY = process.env.OPENAI_API_KEY || 'YOUR_OPENAI_KEY'; const EMBEDDING_MODEL = "text-embedding-ada-002"; const LLM_MODEL = "gpt-4o"; // Or whichever powerful LLM you prefer const LLM_API_URL = "https://api.openai.com/v1/chat/completions"; const EMBEDDING_API_URL = "https://api.openai.com/v1/embeddings"; // --- Step 1: Get Embedding for Query --- async function getQueryEmbedding(query: string): Promise { try { const response = await fetch(EMBEDDING_API_URL, { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': `Bearer ${OPENAI_API_KEY}`, }, body: JSON.stringify({ input: query, model: EMBEDDING_MODEL, }), }); if (!response.ok) { throw new Error(`Embedding API error: ${response.statusText}`); } const data: EmbeddingResponse = await response.json(); return data.data[0].embedding; } catch (error) { console.error("Error getting query embedding:", error); throw error; } } // --- Step 2: Retrieve Relevant Documents from Vector DB --- // This is a conceptual placeholder. In a real system, you'd use a client // for a vector database like Qdrant, Pinecone, Weaviate, or a custom HNSW implementation. // For competitive programming efficiency, I'd likely opt for a highly optimized // local solution like Faiss (with a wrapper) or a custom HNSW tree. async function retrieveDocuments(queryEmbedding: number[], topK: number = 3): Promise { console.log(`Searching vector database for top ${topK} documents...`); // Simulate retrieving documents. In a production environment, this would // query a vector database (e.g., using a client for Agno, Qdrant, etc.) // and return documents whose embeddings are closest to queryEmbedding. // Placeholder data for demonstration const mockDocuments: Document[] = [ { id: "doc1", content: "Mamba models are a new class of deep sequence models based on structured state space models (SSMs)." }, { id: "doc2", content: "Retrieval-Augmented Generation (RAG) improves LLM factual accuracy by fetching relevant external information." }, { id: "doc3", content: "Transformers suffer from quadratic complexity with respect to sequence length due to the self-attention mechanism." }, { id: "doc4", content: "Large Language Models often hallucinate or provide outdated information if not grounded in current data." }, { id: "doc5", content: "State Space Models (SSMs) offer linear time complexity, making them efficient for very long sequences." }, ]; // Simple similarity search (conceptual for demonstration, real one uses vector math) // Here we're just returning some relevant mock docs return mockDocuments .filter(doc => queryEmbedding.length > 0) // Just to use queryEmbedding somehow conceptually .slice(0, topK); // Return topK for simplicity } // --- Step 3 & 4: Construct Prompt and Generate Response --- async function generateResponseWithRAG(userQuery: string, topKDocs: number = 3): Promise { // 1. Get embedding for the user query const queryEmbedding = await getQueryEmbedding(userQuery); // 2. Retrieve relevant documents const relevantDocuments = await retrieveDocuments(queryEmbedding, topKDocs); // 3. Construct the prompt with retrieved context const context = relevantDocuments.map(doc => doc.content).join("\n---\n"); const systemPrompt = `You are a helpful AI assistant. Answer the following question based *only* on the provided context. If the answer cannot be found in the context, state that you don't have enough information.`; const userMessage = `Context:\n${context}\n\nQuestion: ${userQuery}\nAnswer:`; // 4. Call the LLM API with the augmented prompt try { const response = await fetch(LLM_API_URL, { method: 'POST', headers: { 'Content-Type': 'application/json', 'Authorization': `Bearer ${OPENAI_API_KEY}`, }, body: JSON.stringify({ model: LLM_MODEL, messages: [ { role: "system", content: systemPrompt }, { role: "user", content: userMessage }, ], temperature: 0.2, // Keep it low for factual responses }), }); if (!response.ok) { throw new Error(`LLM API error: ${response.statusText}`); } const data: ChatCompletionResponse = await response.json(); return data.choices[0].message.content.trim(); } catch (error) { console.error("Error generating response with RAG:", error); throw error; } } // --- Example Usage --- // (async () => { // const query = "What are the advantages of Mamba models over Transformers?"; // try { // const answer = await generateResponseWithRAG(query); // console.log("\n--- RAG Generated Answer ---"); // console.log(answer); // } catch (e) { // console.error("Failed to get RAG answer:", e); // } // })(); ``` Two lines in there do the actual grounding. The system prompt restricts the answer to the provided context and tells the model to say it doesn't have enough information otherwise, and `temperature: 0.2` keeps the generation close to that context. Everything else is transport. Doing it directly leaves me in control of optimization, caching, and error handling, which matters once you're wiring RAG into a larger system. **The tell:** if you can't point at the line where retrieved text enters the prompt, you don't have RAG. You have a search box next to a chatbot. --- ### Where the two compose The MoE idea I keep coming back to leans on both. Picture experts inside it: some Mamba-like, handling sequential data efficiently, others backed by RAG for specific knowledge domains. A router sends each query to the right experts, and retrieval keeps their knowledge current. ```mermaid flowchart LR Q[query] --> R{router} R -->|sequence task| M[Mamba expert] R -->|knowledge task| G[RAG expert] G --> DB[(vector store)] DB --> G M --> A[answer] G --> A ``` The retrieval edge is the only one that loops, and it's the only one that stays current without retraining anything. --- ### What breaks it **The conceptual scan is not Mamba.** The Python above walks the sequence in the interpreter one timestep at a time, then loops again over the batch. That is exactly what the hardware-aware parallel scan exists to avoid. Read it as a picture of the recurrence, not as an implementation. **A fixed-size state is lossy by construction.** Each step sees the previous state and the current input, and nothing else. If the state didn't carry something forward, the model cannot go back for it. The quadratic cost of attention is what buys random access to every earlier token, and linear scaling is what you pay for by giving that up. **RAG is only as grounded as the retrieval.** The system prompt tells the model to answer only from the provided context. When retrieval returns the wrong documents, that instruction stops being a guardrail and becomes a way to make a wrong answer sound sourced. `retrieveDocuments` above is a placeholder that slices the first `topK` mock documents. The real version of that function is where most of the difficulty lives. --- ### When Mamba and RAG are the wrong choice - **Your sequences are short.** $O(N^2)$ on a small $N$ is a cost you don't have. Swapping architectures to fix arithmetic that never bit you is complexity for nothing. - **The task needs arbitrary reach into the history.** Attention can read any previous token directly. A recurrent state can only read what it decided to keep. If correctness depends on a token far back that selectivity had no reason to retain, the fixed state is the wrong shape. - **You aren't writing kernels.** Mamba's speed comes from careful parameterization plus a hardware-aware scan. Without both, the recurrence you write yourself is a slow loop with extra steps. - **There's no corpus to retrieve from.** RAG grounds in documents. If the knowledge exists only in the weights, retrieval returns nothing useful, and a prompt that says "answer only from the context" turns into a model refusing things it actually knows. - **You don't want the control.** The raw `fetch` version buys optimization, caching, and error handling decisions. If you have no intention of making those decisions, it's just code you now maintain. --- ### The shift Progress here won't come from a single breakthrough model so much as from combining better architectures with better system design. Mamba and RAG are the two halves of that, and they fix different failures. - **Count what one step reads.** Cost scaling is a property of that, not of what the block is named. - **Selectivity is about where the parameters come from.** Input-dependent $A$, $B$, $C$ is the whole difference between a state-space model and a fixed recurrence. - **Put changing knowledge outside the weights.** Weights freeze at the cut-off. Documents don't. Pick by the failure you actually have: cost that grows with length, or knowledge that went stale. --- Source: https://himanshuat.com/blogs/understanding-transformers-part-5-future-architectures ======================================================================== --- title: "Virtual try-on is not image generation" description: "A text-to-image model invents a plausible shirt. Try-on has to reproduce the exact shirt the customer is already looking at, on a person whose face has to survive the process intact. Getting that wrong on a saree is the failure that started all of this." date: "January 14, 2025" url: "https://himanshuat.com/blogs/virtual-try-on-is-not-image-generation" --- # Virtual try-on is not image generation The saree render looked, at a glance, like a saree. The pleats were a printed pattern rather than a set of folds with their own shadows, and the whole thing sat flat on the skin. On the global half of our eval set that pipeline scored SSIM 0.45 against the reference garment. The Western half of the same set scored 0.82. Same model, same code, same week. > A try-on model has exactly one correct output. Every other image it can produce is a bug, however good it looks on its own. I spent about a week sliding between two problems without noticing they were different ones. Generating a picture of someone wearing a shirt is solved. Generating a picture of someone wearing *that* shirt, the one already photographed on white in the brand's catalogue, with the same weave and the same logo placement and the same button spacing, is not. ### Chapter 0 — What a try-on job actually is Two regions, two policies. **Garment fidelity** is a constraint, not a preference. The customer is looking at the garment on the next tab over. They already know what it looks like. Give them a slightly different collar and the render is wrong, and no amount of aesthetic quality fixes it. **Identity preservation** is the other one. The person in the output has to be the same person who was in the input, down to the face, the hands, the hair, the skin tone at the collarbone. Brands ship these images next to a buy button. A model that shifts a face by five percent produces something that looks fine in isolation and deeply wrong next to the original. We had brand teams reject renders for reasons they couldn't articulate beyond "that's not her", and they were always right. Text-to-image respects neither constraint by construction. It has no mechanism to be told "this region is immutable" and no mechanism to be told "this texture is the ground truth, not a hint". Plausible is its whole product, and plausible is the failure mode here. --- ### Step 1 — Write the constraints down before you pick a model I lost time by treating this as a model selection problem when it was a constraint satisfaction problem. Once "the garment is ground truth" and "the person is immutable" were on a whiteboard, the architecture stopped being a choice. One policy for the whole frame cannot satisfy two constraints that point in opposite directions. Everything downstream falls out of those two lines: inpainting over img2img, the mask work, the second conditioning stream. **The tell:** if you can't say in one sentence which pixels are allowed to change, you are not ready to choose a pipeline. --- ### Step 2 — Split the eval set along the axis where you expect bias The open checkpoints everyone starts from, Stable Diffusion 1.5 and SDXL, and the try-on datasets built around them, are overwhelmingly Western casual wear. T-shirts, jeans, hoodies, jackets. Those garments share a topology: they are approximately tubes that sit close to the body, and the body underneath determines the silhouette. A model that has seen a million t-shirts has an excellent prior for fabric conforming to a torso. That prior is wrong for a lot of what our clients sell. A saree isn't determined by the body. It's determined by pleats, by the drape of the pallu over one shoulder, and by gravity acting on several metres of fabric that is not attached to anything. With only the tube prior, the model paints it as a flat texture on skin. The failure gets worse at the semantic level. Ask a stock inpainting model for a kimono and there's a good chance you get a bathrobe. That isn't a rendering bug. The model's nearest concept to "loose belted robe" comes from a distribution where that shape is a bathrobe, so it confidently hands you a bathrobe with nicer fabric. **You cannot prompt your way out of a missing concept.** We called this Western Bias internally, and it stayed a talking point until we had an eval set to put a number on it. Five hundred images, split evenly between Western casual and global traditional. | base pipeline | western casual | global traditional | |---|---|---| | SDXL + ControlNet | 0.82 | 0.45 | | Flux Fill | 0.91 | 0.58 | That spread is the thing worth naming, so we named it the Complexity Gap. Changing the base inpainter moved both halves, and the global half at its best still sits below where the Western half started. A single averaged quality score across that set went up every week while the half of the set our market actually needed barely moved. There is a second axis of the same problem with nothing to do with culture. Complex poses put hands on the garment: a hand on a hip, a hand holding a lapel, arms crossed, a bangle sitting on top of a sleeve. Standard inpainting has no notion of depth ordering inside the masked region, so it paints garment texture straight over the hand. We called that one the Amputated Hand, and it is the most reliable way to make a person recoil from an otherwise good image. **The tell:** if your quality metric is one number, you cannot tell the difference between a system that improved and a system that improved on the half you were already good at. --- ### Step 3 — Use inpainting, not img2img The first thing I tried was the obvious thing. I'm keeping it here because the way it fails points straight at the fix. ```python # tryon_img2img_naive.py """First attempt at try-on: img2img over the model photo with a garment prompt. Kept in the repo because the failure mode is instructive, not because it works. """ import torch from diffusers import StableDiffusionXLImg2ImgPipeline from PIL import Image MODEL_PHOTO = "data/model_0412.png" PROMPT = ( "a woman wearing a deep green silk saree with a gold zari border, " "studio lighting, plain background, full body" ) pipe = StableDiffusionXLImg2ImgPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-base-1.0", torch_dtype=torch.float16, ).to("cuda") def sweep_strength(image_path: str, prompt: str, values: list[float]) -> None: """Run img2img at several denoise strengths and write each result to disk. `strength` is the fraction of the diffusion trajectory we re-noise before denoising back. It is a single global knob applied to every pixel. """ init = Image.open(image_path).convert("RGB").resize((1024, 1024)) for s in values: out = pipe( prompt=prompt, image=init, strength=s, guidance_scale=6.0, num_inference_steps=30, ).images[0] path = f"out/img2img_s{int(s * 100):03d}.png" out.save(path) print(f"strength={s:.2f} -> {path}") sweep_strength(MODEL_PHOTO, PROMPT, [0.35, 0.55, 0.75, 0.95]) ``` Run that sweep and put the four outputs side by side. At low strength the original garment shows through the new one, because you haven't destroyed enough of the original signal to replace it. At high strength you get a convincing saree on a woman who is not the woman you started with: her face has drifted, her arms are a different length, and the background studio seam has moved. No value of `strength` works, and the reason is structural. img2img gives you one policy for the entire frame. Try-on needs two: destroy everything inside the garment region, preserve everything outside it exactly. That is the definition of inpainting. **The tell:** if you are sweeping a global parameter looking for the value that satisfies two opposing constraints, the parameter is the wrong shape for the problem. --- ### Step 4 — Spend your engineering on the mask Once the pipeline is inpainting, the interesting work moves into the mask, because if the mask is wrong nothing downstream saves you. Too tight and you get a halo of the old garment around the collar and cuffs. Too loose and the model starts inventing shoulders. We ended up on SAM2 for the boundary, with a hierarchical prompt set, and a morphological dilation on top. ```python # build_garment_mask.py """Produce the inpainting mask for a single try-on job. SAM2 gives us the garment boundary. Dilation gives the diffusion model a few pixels of adjacent skin and background to blend against, which is what stops the composite from showing a hard seam at the neckline. """ import cv2 import numpy as np from sam2.build_sam import build_sam2 from sam2.sam2_image_predictor import SAM2ImagePredictor GARMENT_PROMPTS = ["shirt", "tshirt", "top", "saree", "dress"] DILATION_PX = 11 # the bleeding zone; we ran anywhere from 5 to 15 MIN_MASK_AREA_FRAC = 0.04 # below this the segmentation has clearly missed predictor = SAM2ImagePredictor( build_sam2("sam2_hiera_l.yaml", "checkpoints/sam2_hiera_large.pt") ) def garment_mask(image_bgr: np.ndarray, prompts: list[str]) -> np.ndarray: """Return a dilated binary mask covering the garment region. Raises if the winning mask is implausibly small, which in practice means the subject is turned away or the garment is mostly out of frame. """ predictor.set_image(cv2.cvtColor(image_bgr, cv2.COLOR_BGR2RGB)) masks, scores, _ = predictor.predict(prompts=prompts, multimask_output=True) best = masks[int(np.argmax(scores))].astype(np.uint8) frac = float(best.sum()) / best.size if frac < MIN_MASK_AREA_FRAC: raise ValueError(f"mask covers {frac:.3f} of the frame, refusing to inpaint") kernel = cv2.getStructuringElement( cv2.MORPH_ELLIPSE, (DILATION_PX, DILATION_PX) ) dilated = cv2.dilate(best, kernel, iterations=1) grown = (float(dilated.sum()) / dilated.size) - frac print(f"mask frac={frac:.3f}, grew by {grown:.3f} at {DILATION_PX}px dilation") return dilated * 255 ``` *The `prompts` argument is our own wrapper, not upstream SAM2's signature; the resolver that turns a garment label into concrete box prompts is glue we wrote, and it isn't the interesting part.* `DILATION_PX` is. That band of dilated pixels is the only place where the model sees both fabric and skin at once, and it is where the transition between them gets decided. Too thin and the neckline reads as a sticker. Too thick and you have handed the model permission to modify the collarbone, which it will take. **The tell:** if your renders show a hard seam at the neckline, you are tuning the wrong stage. That's a mask number, not a sampler number. --- ### Step 5 — Condition on pixels, not on a description of them The mask tells the model where to paint. Something else has to tell it what to paint, and a text prompt is a terrible channel for that. "Green silk saree with gold zari border" is maybe forty bits of information about an object with thousands of bits of relevant structure. That's where the dual-stream setup came from. Flux Fill handles the inpainting. Flux Redux runs alongside it and injects structural features from the reference garment image itself into cross-attention. ```mermaid flowchart LR G[garment ref] --> R[redux encode] M[model photo] --> S[SAM2 mask] R --> F[flux fill] S --> F M --> F F --> P[refine pass] P --> C[4K composite] ``` ```python # tryon_dual_stream.py """Dual-stream try-on: Flux Fill inpaints, Flux Redux conditions on structure. The Redux stream is what separates this from prompt-driven inpainting. A CLIP vision encoder would compress the reference to something like "a red knit sweater"; Redux keeps enough structure to reproduce the knit itself. """ from dataclasses import dataclass @dataclass(frozen=True) class TryOnConfig: """Inference settings for a single garment swap.""" steps_primary: int = 30 steps_refine: int = 10 sampler: str = "euler_ancestral" scheduler: str = "beta" native_resolution: int = 1024 redux_weight: float = 1.0 def run_tryon(model_image, garment_image, mask, cfg: TryOnConfig): """Composite `garment_image` onto `model_image` inside `mask`. Returns the inpainted frame at the model's native resolution. Compositing back into the original 4K frame happens one layer up. """ style = redux_encoder(sigclip(garment_image), weight=cfg.redux_weight) latents = flux_fill( image=model_image, mask=mask, conditioning=style, steps=cfg.steps_primary, sampler=cfg.sampler, scheduler=cfg.scheduler, width=cfg.native_resolution, height=cfg.native_resolution, ) refined = flux_fill( image=latents, mask=mask, conditioning=style, steps=cfg.steps_refine, sampler=cfg.sampler, scheduler=cfg.scheduler, denoise=0.35, ) print(f"primary {cfg.steps_primary} steps, refine {cfg.steps_refine} steps") return refined ``` The named failure this fixes is Style Drift: the output matches the reference in colour and silhouette but loses the material. A cable-knit sweater arrives as a flat red top. A brocade arrives as printed fabric. > Colour survives compression to a semantic vector. Weave does not. If your only conditioning path is semantic, weave is what you lose. **The tell:** if your outputs are the right colour and the wrong material, you have one conditioning stream where you need two. --- ### What breaks it **Sheer fabrics.** Lace and chiffon need the model to blend skin tone *through* the fabric rather than paint over it, and we render them opaque. I don't have a fix that doesn't involve a separate alpha-aware path. **Extreme poses.** Anything outside standing and sitting degrades, because the base model and everything we trained on top of it is biased toward catalogue poses. Yoga and dance shoots fail. This is the same shape of problem as the Amputated Hand: nothing in the pipeline reasons about what is in front of what. **The Complexity Gap is still open.** Getting from 0.58 on global traditional to something a brand signs off on is not a mask problem or a conditioning problem. It's a missing prior, and priors have to be trained. --- ### When constrained try-on is the wrong choice Everything above is machinery for satisfying two constraints simultaneously. If you don't have both constraints, you're paying for it and getting nothing. - **The customer never sees the source garment.** Mood imagery, concept boards, campaign exploration. Nothing is being reproduced, so plausible is a correct answer and text-to-image is the right tool. - **The person is not a specific person.** If identity is not immutable (a synthetic model, a mannequin, a crop with no face in frame), you have one constraint rather than two, and the mask work stops earning its cost. - **Your catalogue is Western casual only.** Most of this post is a response to the global half of our eval set. If your entire line is t-shirts and jackets, the tube prior in the base weights is already doing the hard part for you. - **The garments you care about are sheer, or the shoots are not catalogue poses.** Those are the two places this pipeline still fails on our own eval set. Don't build on top of a known failure and hope. **The tell is Step 2.** Split your set and score both halves. If the halves agree, the base prior covers your catalogue and you're arguing about sampler settings. If they don't, the distance between them is your actual roadmap. --- ### The shift I picked the architecture first and wrote the constraints down second, which is how a week disappeared into an img2img strength sweep that could not have worked. The more expensive mistake was the metric. A single averaged score across the eval set rose every week and told me nothing, because the half it was rising on was the half we already had. Split the metric before you have evidence of bias, not after. - **Say which pixels are immutable, then choose the pipeline.** The constraints pick the architecture, not the other way around. - **Score every half of your eval set separately.** An average across a biased set is a number that agrees with you. - **Hand the model pixels wherever you can hand it pixels.** Text is the weakest structural channel available. Naming the failures did more for our review cycle than any single model change: Western Bias, the Complexity Gap, Style Drift, the Amputated Hand. Once a failure has a name, a brand reviewer can tell you which one they're looking at instead of telling you it feels off. Next is the missing prior. I'm training a LoRA on nothing but garments whose shape comes from gravity rather than from the body underneath, sarees and hanfus and kimonos and haute couture, to see whether drape can be taught as a specialisation instead of baked into a foundation model none of us can afford to retrain. --- Source: https://himanshuat.com/blogs/virtual-try-on-is-not-image-generation ======================================================================== --- title: "We beat the benchmarks and ran out of runway" description: "Flux-VTON+ scores 0.85 SSIM on global traditional garments where SDXL manages 0.45, and we have signed deals with real apparel groups. We're winding the company down anyway. This is what I think the research was worth, written while both of those things are true at once." date: "September 25, 2025" url: "https://himanshuat.com/blogs/we-beat-the-benchmarks-and-ran-out-of-runway" --- # We beat the benchmarks and ran out of runway We started AuraX in December 2024. We're winding it down. The pipeline works. Flux-VTON+ takes SSIM on global traditional garments from 0.45 to 0.85, we were incubated at IIIT-Hyderabad's CIE, and we signed commercial deals with Aditya Birla Fashion Group and several Mensa Brands labels. The decision to stop is made even though the last of the paperwork isn't, so I'm writing this from the middle rather than from the other side. > A benchmark win is evidence about a model, not evidence about a business, and I let myself blur those for most of a year. When our numbers came back strong, I read it as a signal about our position in the market. It wasn't one. It was a signal about our position in a table. ### Chapter 0 — What the complexity gap is Most open-source diffusion models, and the try-on datasets built around them, skew heavily toward Western casual wear. T-shirts, jeans, hoodies: garments with simple topology, close to a tube around the body. A saree is not that. It depends on pleats, on gravity, on the pallu falling over the shoulder in a specific way, and a model without that prior renders it as a flat texture painted onto skin. Standard models classify a kimono as a bathrobe, because that's the nearest thing in their training distribution. So the number I cared about was never a score. It was a difference: **a method's SSIM on Western garments minus its SSIM on global traditional garments, on the same eval set.** We called it the complexity gap. SDXL with ControlNet carries a gap of 0.37. Ours is 0.09. That's a dataset bias problem wearing a computer vision costume. It affects a very large number of people who happen not to be well represented in the images scraped off the Western internet, and the fashion industry's answer for those people has been that the technology doesn't work well yet. --- ### Step 1 — Split the eval set before you trust the average The evaluation set was 500 images, split evenly between what we called Western Casual and Global Traditional. That split was the whole point. Anyone can post a good t-shirt result. Had we reported one averaged SSIM across all 500, we would have posted a fine number, believed our own progress, and shipped something Indian brands would have rejected on sight. It is the single best methodological decision we made, and it cost nothing but the discipline to label the images twice. **The tell:** if your eval reports one average, you don't know which half of the set is carrying it. --- ### Step 2 — Report the gap, not the win Three methods, one eval set, four measurements each. | method | FID | SSIM western | SSIM global | occlusion acc | |---|---|---|---|---| | SDXL + ControlNet | 28.4 | 0.82 | 0.45 | 62% | | Base Flux Fill | 22.1 | 0.91 | 0.58 | 70% | | Flux-VTON+ (ours) | 18.5 | 0.94 | 0.85 | 92% | Occlusion accuracy is human raters judging whether a hand resting on a hip survived the generation intact. On Western garments our SSIM improvement was respectable and unremarkable, because Western garments were never the hard part. The base Flux Fill row is the one that keeps the table honest: it sits between the two on everything, which told us how much of the gain came from choosing a better base model and how much came from the work we added on top. The claim we were proudest of is the success rate on complex global garments, **85% for our pipeline against 15% for the baselines.** That gap is the entire company thesis in one line. ```python # eval/report.py """Tabulate the Flux-VTON+ evaluation run. The eval set is 500 images, split evenly between Western Casual and Global Traditional. Every number here comes out of a scored run; nothing is entered by hand except the method labels. """ from dataclasses import dataclass @dataclass class MethodResult: name: str fid: float # lower is better ssim_western: float # higher is better ssim_global: float # higher is better occlusion_acc: float # human raters, higher is better RESULTS = [ MethodResult("SDXL + ControlNet", 28.4, 0.82, 0.45, 0.62), MethodResult("Base Flux Fill", 22.1, 0.91, 0.58, 0.70), MethodResult("Flux-VTON+ (ours)", 18.5, 0.94, 0.85, 0.92), ] def complexity_gap(result: MethodResult) -> float: """How much worse a method is on global attire than on Western attire. This is the number I actually cared about. A method can post a strong average and still be useless to an Indian apparel brand if the entire average is carried by t-shirts. """ return result.ssim_western - result.ssim_global def report(results: list[MethodResult]) -> None: for r in results: print( f"{r.name:<22} " f"FID {r.fid:>5.1f} " f"SSIM-W {r.ssim_western:.2f} " f"SSIM-G {r.ssim_global:.2f} " f"Occ {r.occlusion_acc:.0%} " f"gap {complexity_gap(r):.2f}" ) if __name__ == "__main__": report(RESULTS) best = min(RESULTS, key=complexity_gap) print(f"\nsmallest complexity gap: {best.name} ({complexity_gap(best):.2f})") ``` `complexity_gap` is the honest summary of the whole research effort. Everything else in the table is a score. That one line is a diagnosis. **The tell:** if you can't compute a spread across the hard and easy halves of your eval, you are reporting a leaderboard position, not a capability. --- ### Step 3 — Close the gap with adapters, not a retrain We closed most of the 0.37 gap with two LoRA experts. One trained on draping physics across sarees, hanfus, kimonos and haute couture. One trained on occlusion and depth. Five thousand curated images of draping. Three thousand of hands on hips and jewellery over fabric. Rank 32, alpha 16, a learning rate of 1e-4 on cosine annealing, a script anyone can run. The cost is a few thousand labelled images per adapter and a single training run, not a foundation model retrain, which costs an amount neither we nor anyone in our position has. The merge is the part worth copying. Both experts go into the backbone at fixed weights rather than being swapped in at runtime. ```mermaid flowchart LR B[Flux Fill] --> M[merge] D[draping LoRA] -->|0.6| M O[occlusion LoRA] -->|0.4| M M --> W[Flux-VTON+] ``` Four lines of arithmetic on top of a public model, and the gap goes from 0.37 to 0.09. **The tell:** if a few thousand curated images move your hard-half score by that much, you had a coverage problem, not a capability problem, and no amount of scaling was going to find it for you. --- ### Step 4 — Name the failure mode in the buyer's vocabulary Naming the failure mode was worth more than solving it. "The Complexity Gap" and "the Amputated Hand" gave us a way to talk to brands about what was broken. A brand that recognises its own problem in your vocabulary listens differently, because you've stopped selling a model and started describing their Tuesday. The names travelled further than the metrics did. Nobody in a procurement meeting repeated 0.85 back to me. Several people repeated the amputated hand. **The tell:** when the customer starts using your phrase for the problem unprompted, the name is doing work no metric on your slide is doing. --- ### Step 5 — Price the thing the benchmark doesn't measure ai.fashion has raised $3.6 million. alphabake has raised $500K. Those are the two figures I had in front of me while writing our technical documentation, and they were sobering in a way that a benchmark table is not. Funding buys things that don't show up in FID. It buys a sales team that can sit through a six-month procurement cycle at a large apparel group without the company dying in month four. It buys the plugin work, the SDKs, the design partner hand-holding, the second and third attempt at positioning. It buys the ability to be wrong for a while. We have signed deals with real names. Aditya Birla Fashion Group is not a logo you put on a slide speculatively, and multiple Mensa Brands labels are working with us. I won't put volumes next to those, partly because I won't publish them and partly because the number that mattered was never in a contract. It's the distance between when we could get revenue at scale and when we run out of room to operate. Larger players entered the space and the runway was short. Those two facts arrived close enough together that arguing about which one is decisive doesn't help anyone. **The tell:** if you can quote your FID to one decimal place and can't quote your nearest competitor's raised capital, you are measuring the half of the problem you enjoy measuring. --- ### What breaks it **A benchmark answers only the question you encoded into it.** Ours encoded: given a model image and a garment image, does the result look right, particularly for garments the field had been ignoring. It answered that well. It said nothing about whether a brand's creative director would approve the output on a Tuesday. Nothing about how many images a catalogue team needs before the tool changes their workflow instead of adding to it. Nothing about integration effort, procurement cycles, or whether the person championing us internally would still be in that role in six months. Nothing about what happens when a well-funded competitor ships something merely adequate with a Shopify plugin already built. **The comparison you grade yourself.** Against the commercial try-on products in our space, Alphabake and FashnAI, our comparison was qualitative: images side by side, no scores. Our own writeup notes that some of the competitor outputs we compared against were generated from lower quality garment and model inputs, which is exactly the caveat you should be suspicious of when a company writes it about a competitor. *I believed our results were better and I still do.* I also know what a fair comparison would have required, and we didn't run it. **Spending your fastest capability where you're already strongest.** I don't think we made one identifiable mistake that ended the company. That would be an easier story to tell. What I think happened is that we spent our strongest asset, research speed, on the problem where we were most capable rather than the problem where we were most exposed, and by the time the exposure was obvious the option to fix it needed capital we didn't have. --- ### When winning the benchmark is the wrong thing to optimise Model quality was not the binding constraint on our outcome. That sentence took me most of this year to be able to write plainly, because every instinct I have as a researcher pushes toward the belief that a better model wins. - **The buying decision isn't made on output quality.** If the gate is a creative director's approval, an integration, or a procurement cycle, a better score moves nothing that the buyer is actually weighing. - **An adequate competitor already ships the integration.** Merely adequate output with a Shopify plugin behind it beats a better model that a catalogue team has to wire up themselves. - **You can't afford the comparison that would settle it.** If your competitive claim is images side by side with no scores, you have a belief, not a result. We had a belief. - **Your exposure is capital and time rather than capability.** Those two are the same variable, and no amount of research speed converts into more of it. --- ### The shift I'd still choose the research problem. Given a year and the same choice, I would work on non-Western garment rendering again, because it was underserved, tractable, and nobody with more resources was doing it seriously. What I'd change is how much of my attention went to the model versus how much went to the years of oxygen a company needs in order to sell one. Last week I put the saree case back through base Flux Fill, without our experts merged in, to check I hadn't been generous to ourselves in the writeup. The pallu still comes out as a flat printed panel stuck to the shoulder, with no fold and no shadow where the fabric gathers at the waist. The gap we named is exactly where we found it. - **Split the eval and report the spread.** One average tells you which half you got lucky on, and it tells you too late. - **Name the failure before you fix it.** The name outlives the metric and travels into rooms you aren't in. - **Spend your speed where you're most exposed, not where you're most capable.** The exposure is obvious only after the option to fix it has expired. The finding doesn't belong to AuraX and it doesn't stop being true because we're stopping. The paper is written, the method is described, and the merge is four lines of arithmetic on top of a public model. If someone with more runway picks it up, I'll take that outcome. > The weights are on a drive in Hyderabad next to a 500-image evaluation set that took months to assemble, and I haven't decided yet what to do with either. --- Source: https://himanshuat.com/blogs/we-beat-the-benchmarks-and-ran-out-of-runway ======================================================================== --- title: "What 5,000 images taught me about curation" description: "More data made the model worse in a way no automatic metric could see. The seed adapters behind AuraX-V1 ended up trained on 5,000 images a person had looked at one by one. Here is what the filter caught, and what it kept missing." date: "February 25, 2025" url: "https://himanshuat.com/blogs/what-5000-images-taught-me-about-curation" --- # What 5,000 images taught me about curation The model wasn't good enough, so I went and got more data. A few weeks later it was worse. Worse in a specific, embarrassing way. It produced images that scored well on everything we could measure automatically, and that our clients would not put on a product page. The fix ran in the other direction. The seed adapters behind AuraX-V1 were trained on 5,000 curated images, and the curation is the part I'd defend in a room full of people who disagree. > Quantity buys you more of whatever is already common. The categories where we needed help were the rare ones, so every additional pile of scraped images scaled the bias rather than the coverage. Fashion imagery on the open internet is abundant and mostly unusable for this. Not because it's low quality in an obvious sense. Plenty of it is beautiful. It's unusable because the distribution is wrong for the job. Editorial fashion photography is dramatic: strong side lighting, deep shadow, heavy grade, sometimes motion blur as a stylistic choice. E-commerce photography is the opposite. Even lighting, the garment fully legible, a background that doesn't compete, a pose that shows the product rather than the photographer. Train on a large pile of the former and ask for the latter, and you get the former with a lighter background. So the question stopped being how many images we could get. It became which images were worth a gradient step. ### Chapter 0 — What "curated" meant here Three gates in series, and an image had to survive all three. Stage one is mechanical: resolution, aspect ratio, sharpness, blown highlights. Stage two is a score for commercial fitness. Stage three is a person looking at the image. ```mermaid flowchart TD A[candidates] --> B{technical gate} B -->|fail| X[dropped] B -->|pass| C{brand score} C -->|fail| X C -->|pass| D{human pass} D -->|fail| X D -->|pass| E[corpus] ``` The ordering is the whole design. Each stage is more expensive than the one before it, so the cheap mechanical checks get first refusal on everything and the human only ever sees survivors. --- ### Step 1 — Reject on physics before you reject on taste The first stage throws out a lot, and none of it is a judgement about whether the photograph is any good. Checks run cheapest first: resolution, then aspect ratio, then sharpness, then clipped highlights. ```python # filter_public_images.py """First-stage technical filter over candidate training images. Nothing here is about taste. These are the checks that decide whether an image is physically capable of teaching the model anything about fabric. Aesthetic and brand judgement happen downstream, on what survives. """ from collections import Counter from dataclasses import dataclass from pathlib import Path import cv2 import numpy as np MIN_SHORT_SIDE = 768 # below this, weave detail is already gone SHARPNESS_FLOOR = 120.0 # variance of Laplacian on the luma channel ASPECT_RANGE = (0.55, 1.45) # portrait-ish through square; rejects banners MAX_CLIP_FRAC = 0.02 # fraction of pixels blown to pure white @dataclass(frozen=True) class Verdict: """Outcome of the technical gate for one candidate image.""" path: Path keep: bool reason: str def inspect(path: Path) -> Verdict: """Apply the technical gate to a single image. Sharpness uses the variance of the Laplacian, which is a blur proxy, not a quality score. It reliably catches upscaled thumbnails and camera shake and says nothing at all about whether the photograph is any good. """ img = cv2.imread(str(path)) if img is None: return Verdict(path, False, "unreadable") h, w = img.shape[:2] if min(h, w) < MIN_SHORT_SIDE: return Verdict(path, False, "resolution") aspect = w / h if not ASPECT_RANGE[0] <= aspect <= ASPECT_RANGE[1]: return Verdict(path, False, "aspect") gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) if cv2.Laplacian(gray, cv2.CV_64F).var() < SHARPNESS_FLOOR: return Verdict(path, False, "sharpness") if float((gray >= 250).sum()) / gray.size > MAX_CLIP_FRAC: return Verdict(path, False, "blown_highlights") return Verdict(path, True, "ok") def run(root: Path) -> Counter: """Walk a candidate directory and tally why images were dropped.""" tally: Counter = Counter() for path in sorted(root.rglob("*.jpg")): verdict = inspect(path) tally[verdict.reason] += 1 if verdict.keep: print(path) for reason, n in tally.most_common(): print(f"{reason:>18}: {n}", file=__import__("sys").stderr) return tally ``` `ASPECT_RANGE` looks like the least interesting line in that file and turned out to be one of the most useful. A wide crop of a fashion image is almost always a banner, a lookbook spread with three models in it, or a detail shot with no garment context. **It's a cheap proxy for "one person wearing one outfit".** `MAX_CLIP_FRAC` went in after we noticed the model producing garments with flat white regions where a fold should be. Overexposed training images teach the model that white is a legitimate value for a shadow-facing surface, and once it learns that, silk and satin come out looking like paper. `SHARPNESS_FLOOR` is the one people over-trust. Variance of Laplacian catches blur. It cannot tell you an image is well composed or well lit, and if you treat it as a quality score you'll keep a lot of sharp, ugly photographs. **The tell:** if your model paints silk like paper, look at the highlights in your training set before you touch the sampler. --- ### Step 2 — Organise the corpus by the behaviour, not the product category Past the technical gate, the corpus was organised by material. That's the decision I'd most want to explain to someone starting this. The seed adapters had to cover denim, silk and leather. Those aren't product categories, they're behaviours. | material | what it has to hold | how a fake gives itself away | |---|---|---| | denim | a crease, plus a visible twill at zoom scale | the weave dissolves when a viewer zooms in | | silk | specular highlights that move with the drape | the falloff is a painted gradient, not a surface | | leather | a sheen that sits between the other two | the sheen is uniform | Organising this way made the balance question answerable. A jacket in leather and a jacket in denim are two different training examples for our purposes, even though a product catalogue would file them together. ```yaml # seed_corpus.yaml # Manifest for the 5,000-image curated corpus behind the seed adapters. # Every entry below is a gate or a policy. Per-material counts live in the # generated index, not here, because they moved every time we re-reviewed. technical_gate: min_short_side: 768 sharpness_floor: 120.0 aspect_range: [0.55, 1.45] max_clip_frac: 0.02 materials: - denim - silk - leather policy: balance_by: material max_share_single_source: 0.15 # no single origin dominates a material require_full_garment_visible: true reject_if_face_occluded: false # faces are not what these adapters learn duplicate_check: perceptual_hash review: stage_1: technical_gate stage_2: brand_aesthetic_score stage_3: human_pass # every surviving image seen by a person ``` `max_share_single_source` exists because our first balanced corpus was balanced by material and completely unbalanced by photographer. One source can supply a thousand technically excellent silk images that all share a lighting setup, and the adapter learns the lighting setup. Worth keeping straight, because I've conflated it myself in conversation: this corpus is not the same corpus as the other experts. | corpus | images | contents | |---|---|---| | seed adapters | 5,000 | denim, silk, leather | | draping physics | 5,000 | sarees, hanfus, kimonos, complex haute couture | | occlusion and depth | 3,000 | hands on hips, arms crossed, jewellery over clothing | Same order of magnitude, unrelated contents, different organising axis. Material was right for the seed adapters. It would be the wrong axis for the occlusion expert, where the axis is the pose. **The tell:** if the same garment in two materials lands in one bucket, you're balancing a catalogue rather than a corpus. --- ### Step 3 — Train the scorer on your customers' taste, not the internet's We tried to automate stage two with an off-the-shelf aesthetic scorer. This is where the interesting failure lives. Generic aesthetic models, including some very good recent ones, have a learned bias toward darker and moodier imagery. Ask one to rank a batch of fashion photographs and it will consistently prefer the one with crushed blacks, a warm grade and a strong key light. That preference is real, and it reflects what people upvote on the internet, which is what those models were trained on. It is also precisely backwards for our customer. No e-commerce brand ships a moody product shot where the garment is half in shadow. The whole job of the image is to show the product clearly. Using a generic scorer as a filter meant systematically selecting training data that pushed the model away from what our clients would accept, while the score went up. > A generic quality metric encodes somebody's preferences. If you don't know whose, you're optimising toward a stranger. So we trained our own on our own labels. A small head over frozen image embeddings, fit on pairwise preferences rather than absolute ratings, because people are bad at giving a photo a 7 out of 10 and quite good at saying which of two photos they'd rather put on a product page. ```python # train_brand_aesthetic.py """Fit the Brand-Centric Aesthetic Model on in-house pairwise preferences. Absolute ratings drift between sessions and between people. Pairwise choices are stable enough to learn from, so every label is "A over B for a product page", and the head learns a scalar whose ordering matches those choices. """ import torch import torch.nn as nn from torch.utils.data import DataLoader, Dataset EMBED_DIM = 1024 HIDDEN = 256 LR = 1e-4 EPOCHS = 20 MARGIN = 0.2 class PreferencePairs(Dataset): """Precomputed embeddings for (preferred, rejected) image pairs.""" def __init__(self, pairs: list[tuple[torch.Tensor, torch.Tensor]]): self.pairs = pairs def __len__(self) -> int: return len(self.pairs) def __getitem__(self, i: int): return self.pairs[i] class AestheticHead(nn.Module): """Maps a frozen image embedding to a single commercial-fitness scalar.""" def __init__(self, dim: int = EMBED_DIM, hidden: int = HIDDEN): super().__init__() self.net = nn.Sequential( nn.Linear(dim, hidden), nn.GELU(), nn.Dropout(0.1), nn.Linear(hidden, hidden // 2), nn.GELU(), nn.Linear(hidden // 2, 1), ) def forward(self, x: torch.Tensor) -> torch.Tensor: return self.net(x).squeeze(-1) def train(pairs: PreferencePairs) -> AestheticHead: """Train with a margin ranking loss on preferred-over-rejected pairs.""" head = AestheticHead().cuda() opt = torch.optim.AdamW(head.parameters(), lr=LR) loader = DataLoader(pairs, batch_size=64, shuffle=True) target = torch.ones(1).cuda() criterion = nn.MarginRankingLoss(margin=MARGIN) for epoch in range(EPOCHS): running = 0.0 for better, worse in loader: s_better = head(better.cuda()) s_worse = head(worse.cuda()) loss = criterion(s_better, s_worse, target.expand_as(s_better)) opt.zero_grad() loss.backward() opt.step() running += loss.item() print(f"epoch {epoch:02d} mean_loss={running / max(len(loader), 1):.4f}") return head ``` The honest description of where those pairs came from is that we scored images ourselves and argued about it. *There was no panel, no protocol document, no measured agreement statistic.* It was a handful of people who had spent the previous month on calls with brand teams, sitting with a folder of images and disagreeing out loud until a rule emerged we could apply consistently. Some of those arguments took an hour and produced one line of guidance, like whether a visible garment wrinkle is realism or a defect. (It depends on the material, which is why the corpus is organised by material.) I'm not going to dress that up as a study. It's a small group encoding a commercial taste they'd absorbed from customers, and its main virtue is that it was our customers' taste rather than the internet's. **The tell:** if your scorer's top-ranked batch is the one with crushed blacks and a strong key light, it's ranking for the internet. --- ### Step 4 — Let human attention set the size of the corpus The line I'd underline in that manifest is `stage_3`. Every image that made it into the 5,000 was looked at by a person. That is feasible at five thousand and it is not feasible at fifty thousand, and I think the constraint was doing more work than the number. **Five thousand was never a target we chose.** It's the count that survived a process where somebody looked at every image. Any process that scales past what you can personally inspect had better have something else keeping it honest. **The tell:** if you can't say you looked at every image in the corpus, something other than attention is setting your dataset size, and you should be able to name what. --- ### What breaks it **The scorer learns your reviewing habits.** Early versions learned that images we had cropped in a particular way were preferred, because the ones we'd bothered to crop were the ones we already liked. The label carried our workflow, not our taste. **Balance on one axis is silence on every other.** Balancing by material gave us nothing on body type or skin tone, and I didn't catch that until later than I should have. A corpus can be carefully curated along the axis you're thinking about and unexamined along all the rest. **Curation has a ceiling.** It gets you a model that reliably produces the kind of image you selected for. It cannot produce a kind of image that wasn't in the corpus at all. When a client asked for a look outside what we'd curated, the answer was more curation, not more sampling. --- ### When curation is the wrong choice Curating by hand costs weeks of attention from people who could be building. Several situations don't repay it. - **The base model already covers your categories.** Ours handled t-shirts and denim staples fine. Curation earned its keep only on the categories that were rare in any pile you could assemble quickly. - **You need a look that isn't in the corpus yet.** The ceiling above is a hard one. If the requirement keeps moving, you're signing up for a curation pass per request, not a dataset. - **Your customers' taste is the internet's taste.** The whole argument for training our own scorer was that generic ones point away from e-commerce. If darker and moodier is what your buyers want, an off-the-shelf scorer is already aimed correctly and costs nothing. - **You need more images than a person can look at.** Stage three is what made the other two trustworthy. Drop it and you have an automated filter with nobody checking whether it's rejecting the right things. --- ### The shift Filter for physical capability first and taste second, and keep the two stages separate in code. Mixing them produces a scorer that quietly rejects good photographs for being slightly soft and accepts sharp ones nobody would ship. The darker-and-moodier bias is the clearest example I've hit of a metric working correctly and pointing the wrong way. It wasn't broken. It was answering a question about the internet while we were asking a question about product pages. - **Reject on physics before you reject on taste.** Two stages, two files. - **Organise by what the model has to learn.** Material for fabric behaviour, pose for occlusion. - **Size the corpus to the attention you actually have.** Five thousand is what a person can see. What comes next is closing the loop with the product. Users favouriting and downloading generated images is a preference signal collected at a scale no folder review can reach, and it comes from the people whose taste actually matters. > Turning that into training pairs, without letting it drag the scorer back toward whatever is merely popular, is the problem I'm sitting with now. --- Source: https://himanshuat.com/blogs/what-5000-images-taught-me-about-curation ======================================================================== --- title: "What happens when you merge two LoRAs" description: "Two adapters that each pass on their own can make each other worse the moment you sum them into the base weights. We shipped a hand-tuned coefficient pair because we had to ship something, and months later I still don't have a principled answer for how those coefficients should be chosen." date: "July 15, 2025" url: "https://himanshuat.com/blogs/what-happens-when-you-merge-two-loras" --- # What happens when you merge two LoRAs By June we had two adapters that worked. Expert A put volume back into a saree's pleats. Expert B stopped the pipeline from painting fabric over somebody's hand. Loaded one at a time against the base model, each did the thing it was trained to do, and I had the eval renders to prove it. Then I merged them and watched a kimono come out looking like a bathrobe again. > Two adapters that are individually correct are not jointly correct. The sum is a third model that nobody trained and nobody evaluated. Everything below is an attempt to make that third model shippable, and an honest account of how far short of principled the attempt stopped. ### Chapter 0 — What the merge actually is A LoRA is a low rank update to a weight matrix. You train two small matrices, `B` and `A`, and the adapter's contribution to a layer is their product. Applying an adapter means adding that product to the base weight. Applying two means adding both, with a scalar on each: ```text W_merged = W_base + λ_drape · (B_d A_d) + λ_occ · (B_o A_o) ``` This is the form in our paper and it's the form nearly everyone uses, for a good reason: it's the only combination that costs nothing at inference. You do the arithmetic once, you get a single set of weights, and the model you serve has no idea it was ever two adapters. **The merge is a build step.** The thing that makes it look safe is that it's linear, so it looks like it should compose. Two independent corrections, added independently. If the two updates wrote into disjoint parts of the weight space, it more or less would. They don't. --- ### Step 1 — Print the cosine similarity before you tune anything The whole merge is fifty lines. That's why the operation looks harmless: nothing in it knows the two adapters were trained for different reasons, or that they might be asking the same layer for opposite things. ```python # merge_experts.py """Merge two LoRA adapters into the base weights with per-adapter coefficients. This is the whole merge. It is fifty lines and it is why the operation looks harmless: nothing here knows that the two adapters were trained for different reasons, or that they might be asking the same layer for opposite things. """ from pathlib import Path import torch from safetensors.torch import load_file, save_file def delta_for(lora: dict[str, torch.Tensor], key: str, alpha: float, rank: int): """Reconstruct the dense update BA for one layer, scaled by alpha / rank. Returns None when the adapter did not touch this layer, which is common: a LoRA usually targets the attention projections and leaves the rest alone. """ up = lora.get(f"{key}.lora_up.weight") down = lora.get(f"{key}.lora_down.weight") if up is None or down is None: return None return (up @ down) * (alpha / rank) def merge(base_path: Path, drape_path: Path, occ_path: Path, out_path: Path, l_drape: float = 0.6, l_occ: float = 0.4): base = load_file(base_path) drape = load_file(drape_path) occ = load_file(occ_path) touched_by_both = 0 merged = {} for key, weight in base.items(): stem = key.removesuffix(".weight") d = delta_for(drape, stem, alpha=16, rank=32) o = delta_for(occ, stem, alpha=16, rank=32) if d is not None and o is not None: touched_by_both += 1 # Cosine similarity between the two updates, flattened. Near zero # means they are writing in roughly orthogonal directions and the sum # is close to harmless. Large magnitude, either sign, means they are # fighting over the same subspace. cos = torch.nn.functional.cosine_similarity( d.flatten().float(), o.flatten().float(), dim=0 ).item() if abs(cos) > 0.25: print(f"{stem}: cos={cos:+.3f}") update = torch.zeros_like(weight, dtype=torch.float32) if d is not None: update += l_drape * d.float() if o is not None: update += l_occ * o.float() merged[key] = (weight.float() + update).to(weight.dtype) print(f"layers touched by both adapters: {touched_by_both}") save_file(merged, out_path) if __name__ == "__main__": merge( Path("weights/flux/base.safetensors"), Path("weights/lora/expert_a_drape.safetensors"), Path("weights/lora/expert_b_occ.safetensors"), Path("weights/lora/aurax_merged.safetensors"), l_drape=0.6, l_occ=0.4, ) ``` I added the cosine similarity print to find out how bad the overlap was. The answer is that both adapters target the attention projections, because that's where LoRA training targets by default. They overlap almost everywhere they exist. **The tell:** if `touched_by_both` is large and the printed cosines are far from zero, your adapters are fighting over the same subspace, and no coefficient search is going to fix that. --- ### Step 2 — Separate interference from forgetting Two things are happening. They call for different fixes and I conflated them for weeks. The first is **feature interference**. Both adapters have opinions about the same attention layers. Expert A learned to push the model toward generating structured folds inside the masked region. Expert B learned to push it toward respecting a foreground boundary and generating nothing across it. Summed into one matrix, those pushes partially cancel and partially compound in directions neither adapter was ever evaluated in. The second is **catastrophic forgetting**, and it's about the base. Each adapter individually is a small perturbation. Two of them at full strength is a larger one, and past some magnitude you're not adapting the base model's garment prior anymore, you're overwriting it. That's why the kimono went back to being a bathrobe. That defect isn't a drape failure or an occlusion failure. It's the base model's semantic prior degrading under a perturbation it wasn't built to absorb. You can tell them apart by where the damage lands. | failure | what degrades | where you see it | |---|---|---| | feature interference | the behaviours the adapters were trained for | pleats, folds, occlusion boundaries | | catastrophic forgetting | general competence of the base | hands, faces, backgrounds, garment class | **The tell:** if the things getting worse have nothing to do with either adapter's job, you're overwriting the base, not blending two experts. --- ### Step 3 — Expect style drift to be the first thing you see Before you've worked out which of the two you're looking at, the observable symptom is Style Drift. The generated garment keeps the reference's colour and loses everything else. A knit sweater comes out as a flat red shirt. A brocade saree keeps the gold and loses the weave. The silhouette is right, the palette is right, and the fabric is wrong in a way a brand's art director spots in under a second and a similarity metric mostly doesn't. Drift is what made this expensive rather than annoying. **It's the failure mode that survives your automated checks.** A merged model can look fine across an eval sweep and still be unshippable, because "the fabric doesn't look like the fabric" is a judgement that lives in a human's head. That's a large part of why we ended up building a brand-centric aesthetic scorer with a person in the loop at all. Generic aesthetic scorers pull toward darker, moodier imagery rather than toward commercial e-commerce standards. **The tell:** if your eval sweep is green and the client's art director rejects the batch, the thing you're missing is drift, and you need a human rung in the ladder. --- ### Step 4 — Run the sweep, and be honest about what ends it We landed on `λ_drape = 0.6` and `λ_occ = 0.4`. The paper says empirical testing determined this ratio provides the optimal balance. That's true in the narrow sense that we tested a grid and picked the cell whose renders we liked most. It is not true in any sense that would let you derive it. Here's what the search actually was. ```python # sweep_lambdas.py """Render a fixed prompt set across a grid of merge coefficients. There is no scalar objective at the end of this. It produces a contact sheet and a person decides. That is the honest description of how 0.6 and 0.4 were chosen. """ import itertools from pathlib import Path import torch from diffusers import FluxFillPipeline from merge_experts import merge GRID = [0.2, 0.4, 0.6, 0.8, 1.0] # Four probes, each targeting a behaviour one of the adapters owns, plus one that # neither owns so we can see the base prior degrading. PROBES = [ ("saree_pleats", "full length saree, pleats gathered at the waist, pallu over shoulder"), ("hand_on_hip", "model standing, one hand resting on hip, fitted top"), ("brocade_weave", "heavy brocade fabric, visible gold thread, close weave"), ("plain_tshirt", "plain cotton t-shirt, studio lighting, front facing"), ] def render_all(out_root: Path): for l_drape, l_occ in itertools.product(GRID, GRID): tag = f"d{l_drape}_o{l_occ}" weights = Path(f"tmp/merged_{tag}.safetensors") merge( Path("weights/flux/base.safetensors"), Path("weights/lora/expert_a_drape.safetensors"), Path("weights/lora/expert_b_occ.safetensors"), weights, l_drape=l_drape, l_occ=l_occ, ) pipe = FluxFillPipeline.from_pretrained( "weights/flux/fill", torch_dtype=torch.bfloat16 ).to("cuda") pipe.load_lora_weights(weights) for name, prompt in PROBES: image = pipe(prompt=prompt, num_inference_steps=30, height=1024, width=1024).images[0] path = out_root / tag / f"{name}.png" path.parent.mkdir(parents=True, exist_ok=True) image.save(path) print(f"{tag}: {len(PROBES)} probes rendered") del pipe torch.cuda.empty_cache() if __name__ == "__main__": render_all(Path("sweeps/lambda_grid")) print("contact sheet ready, go look at it") ``` The fourth probe is the one that earns its place. `plain_tshirt` belongs to neither adapter, so it's the cell where you watch the base prior degrade independently of whether the pleats look good. Two things about that coefficient pair bother me and both are still true. It's one scalar per adapter applied to every layer the adapter touches. There is no reason to believe the right weighting for an early attention block is the same as for a late one, and every reason to suspect it isn't, since the layers are doing different jobs. *We're compressing a per-layer question into one number because one number is what the merge API takes.* And the pair is tied to this specific pair of adapters. Train a third expert, for lighting or for a demographic, and the search restarts. **The tell:** if the last step of your tuning process is a person looking at a contact sheet, you have a preference, not an optimum. Say so in the paper. --- ### Step 5 — Pick the composition strategy on serving grounds There are three ways to combine two adapters and I ran two of them. **Keep them separate and swap at runtime.** Load the base once, apply whichever adapter the request needs. No interference, because there's never more than one adapter in the weights. This is genuinely correct and we ran it during development. It falls apart at serving time for two reasons. Applying and unapplying adapters against a resident pipeline is not free. More importantly it forces you to route, which means deciding per request whether this image is a drape problem or an occlusion problem. ```mermaid flowchart LR R[request] --> C{which defect?} C -->|drape| A[expert A] C -->|occlusion| B[expert B] C -->|both| X[no branch] ``` A saree worn by someone with their arms crossed is both. There's no branch to take. **Apply them sequentially.** This sounds like it should help and mostly doesn't. Sequential merging, adding one delta and then the other, is algebraically the same sum you started with; addition doesn't care about order, so nothing changes. Sequential inference is worse. One pass with the drape adapter, a second pass with the occlusion adapter over the first pass's output, and you've doubled your inference cost while the second pass denoises an image the first already committed to. In practice the second pass repairs the hand and softens the pleats the first pass just built. You end up where you started, having spent twice as long. **Fuse once and serve one model.** This is what we needed and what we did. One set of weights, one process, one warm pipeline, no per-request decisions, no VRAM penalty from holding several adapters resident, and a merge step that runs at build time where its cost doesn't matter. | strategy | interference | inference cost | needs routing | |---|---|---|---| | swap at runtime | none | apply/unapply per request | yes, and some requests are both | | sequential inference | none in the weights | doubled | no | | fuse once | yes, this whole post | unchanged | no | In ComfyUI the fuse is a couple of `model_merge_lora` nodes ahead of everything else in the graph. ```json { "31": { "class_type": "LoraLoaderModelOnly", "_meta": { "title": "Expert A: draping physics" }, "inputs": { "lora_name": "expert_a_drape.safetensors", "strength_model": 0.6, "model": ["12", 0] } }, "32": { "class_type": "LoraLoaderModelOnly", "_meta": { "title": "Expert B: occlusion and depth" }, "inputs": { "lora_name": "expert_b_occ.safetensors", "strength_model": 0.4, "model": ["31", 0] } }, "33": { "class_type": "CheckpointSave", "_meta": { "title": "write fused runtime weights" }, "inputs": { "filename_prefix": "aurax/merged_d06_o04", "model": ["32", 0], "clip": ["12", 1], "vae": ["12", 2] } } } ``` The serving argument won, and I'd make the same call again. What I want to be clear about is that it won on operational grounds, not because the merge is correct. **We picked the option whose failure mode we could live with.** **The tell:** if you can't name a request that needs both adapters at once, you don't need to merge. Route instead and skip everything above. --- ### What breaks it **The thing you ship was never evaluated.** Expert A passed its eval. Expert B passed its eval. The sum passed neither, because the sum is a model neither eval was written for. A green sweep on the merged weights measures the behaviours you thought to probe, and drift is specifically the failure that isn't one of them. **The coefficient pair doesn't generalise past this pair of adapters.** The cost of adding an expert grows with the number of experts you already have, which is precisely the property you don't want in a system whose whole premise is that you can patch knowledge gaps with cheap adapters. Our plan for demographics and poses and lighting and background styles was more adapters. That plan and this method don't fit together. **The global scalar hides a per-layer question.** One number for every layer an adapter touches, when the layers are doing different jobs. I don't have a per-layer treatment. I find out whether two new adapters interfere by merging them and rendering probes, which is a measurement, not a prediction. *The thread we've been pulling on since is a composition step that looks at the adapters before they're summed and does something smarter than trusting a human's prior about which one should win.* It works well enough now that it's what our scene generation runs on. It's also changing week to week, and the version I'd describe today isn't the version that would be running by the time you read it, so I'm not going to write up its internals yet. What I can say is that the problem in this post didn't get solved by finding better values for lambda. It got solved by stopping treating lambda as the thing to solve. --- ### When merging two LoRAs is the wrong choice The merged model ships. It's better than either baseline we measured against and it's what our clients' images went through. That's the result and I'm not going to undersell it. It is also the wrong move in at least four situations. - **Your adapters overlap hard.** If the layer-by-layer cosines are far from zero, no value of lambda saves you. You're not blending two experts, you're picking which one loses. - **Every request has one obvious owner.** If a drape problem is never an occlusion problem, routing is strictly better. One adapter in the weights at a time means zero interference and nothing to tune. - **You're planning a fifth adapter.** The search restarts every time you add an expert, so the method gets more expensive exactly as your architecture bets on it getting cheaper. I don't have a story for the fifth adapter, and that's the one that matters. - **Nobody is going to look at the output.** Drift survives automated checks. If there's no human rung anywhere in your pipeline, a merged model will pass your metrics and fail your customer, and you won't find out from the sweep. --- ### The shift Before you spend a week on a coefficient sweep, print the cosine similarity between your two updates layer by layer. If they're near orthogonal, the sum will mostly work and your problem is somewhere else. If they aren't, you'll have learned that in an afternoon instead of a month. - **Measure the overlap before you tune the coefficients.** The cosine print costs one pass over the weights. - **Diagnose which failure you have.** Adapter behaviours weakening is interference. Hands and faces getting worse is forgetting. - **Choose the composition strategy on serving constraints, and say that's why.** Ours won on operations, not on correctness. I shipped a constant somebody chose by looking at pictures, and it's still in production. The honest version of that sentence is the paper I'd rather have written. --- Source: https://himanshuat.com/blogs/what-happens-when-you-merge-two-loras ======================================================================== --- title: "When the documents disagree" description: "An answer cited a specification section by number and was wrong, because an addendum had already revised that clause and nothing in the index knew it. Chasing that led to a FinanceBench result where giving each document its own vector store moved GPT-4-Turbo from 19% to 50% with no change to the model." date: "March 17, 2026" url: "https://himanshuat.com/blogs/when-the-documents-disagree" --- # When the documents disagree The answer was correct, cited a specification section by number, and was wrong. An addendum issued weeks earlier had revised that clause. Nothing in the index knew that one document supersedes part of another, so the citation attached to a dead requirement and made the answer look checked. Someone reads an answer like that and orders material against it. The cost of being wrong is measured in rework and, if it runs far enough, in who pays for it. > A citation is not evidence. It is a claim about evidence, and an index with no model of supersession will produce a true citation to a clause that no longer governs. ### Chapter 0 — What an index actually knows An index is three decisions, and only one of them gets discussed. The first is what a retrieval unit is: a fixed-size chunk, a page range, a section. The second is what scope that unit is searched in: one store over everything, or one store per document. The third is what relations exist between units, which in most systems is none at all. Model choice and prompt phrasing sit outside all three. That is why the rest of this post barely mentions them. --- ### Step 1 — Draw the index boundary before you touch the model The result that explained my failure has nothing to do with construction. The FinanceBench paper by Islam et al., from Patronus AI, Contextual AI and Stanford, works over 10,231 questions across 361 filings from 40 companies, with a 150-question human-evaluated sample. On that sample, the only thing varied between the first two rows below is where the index boundary was drawn. | condition | correct | incorrect | failed to answer | |---|---|---|---| | shared vector store over the corpus | 29 (19%) | 20 (13%) | 101 (68%) | | one store per document | 75 (50%) | 17 (11%) | 58 (39%) | | whole document in context | 118 (79%) | 26 (17%) | not reported here | | oracle: exact evidence pages given | 85% ceiling | 22 (15%) | not reported here | The shared store was Chroma through LangChain on ada-002 embeddings, queried by GPT-4-Turbo. Give each document its own store and the same model on the same questions more than doubles its accuracy. Nothing about the model changed. I keep returning to that pair of rows because of how much attention goes to model selection and prompt phrasing, both of which were held constant across the entire effect. **The tell:** if you can name your embedding model and your reranker but not how many stores your corpus is split across, you have been tuning the layer that moved least. --- ### Step 2 — Find the ceiling before you believe any number Two rows in that table deserve as much attention as the headline, and neither gets quoted. The oracle condition is the first. Hand GPT-4-Turbo the exact evidence pages, removing retrieval from the problem entirely, and it still gets 22 of 150 wrong. **The ceiling is 85%, not 100%.** Any claimed accuracy above that on this benchmark is describing a system where something other than retrieval also improved. Long context is the second, and it is subtler. Feeding the whole document in gives 79% correct against 17% incorrect. Compare that 17% to the shared store's 13% and you find long context did not remove errors so much as convert refusals into wrong answers. The failures stopped announcing themselves. Prompt order matters enormously in that setting too: 78% versus 25% for GPT-4-Turbo depending on whether the context precedes or follows the question. Any long-context comparison that doesn't control for it is measuring the wrong thing. The authors put the asymmetry plainly: "Models refusing to answer is arguably preferable to giving an incorrect answer as it creates less risk of error, and misplaced trust, by users." > A refusal is a cheap failure. A confident wrong answer is an expensive one, and a system with no refusal path can only fail the expensive way. That sentence is the design principle for everything after it. The output type I want is `answer | clarification_request | refusal`, and the refusal arm has to fire often enough to be a real path rather than a branch that exists only in the type signature. **The tell:** if your refusal rate went down after a change and you called that an improvement, check whether those refusals became correct answers or wrong ones. --- ### Step 3 — Navigate the hierarchy the author already wrote The most interesting alternative I've worked with is PageIndex, which deserves a careful reading rather than a credulous one. Its README describes two steps: "(1) Generate a 'Table-of-Contents' tree structure index of documents; (2) Perform (agentic) reasoning-based retrieval through tree search." There is no vector store anywhere in that description. Ingestion runs an LLM pass over the document to build a tree mirroring its table of contents. Retrieval hands the model that tree of titles and summaries, lets it select node IDs, and fetches only those page ranges. The node schema is where the value sits. ```python # node.py from dataclasses import dataclass @dataclass class Node: """A PageIndex-style node. The two index fields are page numbers, which is the property that makes every retrieval result citable by construction.""" title: str # "03 30 00 Cast-In-Place Concrete" node_id: str # stable id the model selects during tree descent start_index: int # first page end_index: int # last page summary: str # what the model reads while navigating children: list["Node"] def descend(root: Node, question: str, budget: int = 4) -> list[Node]: """Walk the tree by reasoning rather than by distance. Each level costs one sequential LLM call, and the calls cannot be parallelised because the next level depends on the selection made at this one.""" selected, frontier, depth = [], [root], 0 while frontier and depth < budget: menu = [ {"node_id": n.node_id, "title": n.title, "summary": n.summary} for node in frontier for n in node.children ] if not menu: break chosen_ids = ask_model_to_select(question, menu) # returns node_ids frontier = [n for n in all_nodes(root) if n.node_id in chosen_ids] selected.extend(frontier) depth += 1 return selected def cite(node: Node) -> str: """No post-hoc citation extraction is needed: the title and page range came back with the content, because they are what the retrieval selected on.""" return f"{node.title}, pp. {node.start_index}-{node.end_index}" ``` Every result carries a section title and a page range because those are the fields retrieval navigated by. The usual approach retrieves a chunk and then reverse-engineers which page and heading it came from, which is a lossy reconstruction of information chunking already discarded. The comparison most write-ups skip is RAPTOR (arXiv 2401.18059, ICLR 2024), which also builds a tree over documents. *What follows is my reading and not an established result: neither project makes this comparison, and I am describing RAPTOR's mechanism from secondary reading rather than from a careful pass over the paper.* RAPTOR is described as working bottom-up. Embed the chunks, cluster them with a Gaussian mixture model, summarise each cluster, repeat over the summaries, then retrieve by embedding similarity across all levels. PageIndex works top-down, taking the document's own table of contents as the structure and navigating it by reasoning. The direction of construction is the whole difference. RAPTOR's hierarchy is induced, its clusters are statistical artifacts, and no node in that tree corresponds to anything a human would cite. PageIndex's hierarchy is the author's own, which is exactly why its nodes are citable. That matters most where the document already ships its structure. As it is commonly described, and *I have this from summaries rather than from the standard itself*, CSI MasterFormat organises construction specifications into 50 divisions, expanded from 16 in November 2004, with six-digit numbering in a Division-LevelTwo-LevelThree pattern. Section 03 11 00 is Division 03 Concrete, subgroup 03 11 Concrete Forming and Accessories. Inside a section, SectionFormat imposes a fixed three-part structure: Part 1 General, Part 2 Products, Part 3 Execution. For a specification book where CSI section numbering is the ground-truth hierarchy, running a clustering algorithm over the text throws away the answer key that came printed in the document. None of this removes approximation. Vector search approximates relevance through embedding geometry; a tree walk approximates it through LLM reasoning over summaries. Both can be wrong, and swapping one for the other moves the approximation rather than deleting it. What changes is inspectability: when a vector search returns the wrong chunk the explanation is a distance in a space you cannot read, and when a tree walk returns the wrong section you can read the path it took, see which summary misled it, and fix that summary. That is a genuine advantage, and a smaller claim than the one usually made on this technology's behalf. **The tell:** ask what a retrieval result would look like printed in an RFI. If the answer is a bare chunk index rather than a section number and a page range, citation is being reconstructed rather than retrieved. --- ### Step 4 — Check the headline claim against the ceiling The number attached to PageIndex is that "Mafin 2.5, powered by PageIndex, achieved a state-of-the-art 98.7% accuracy on FinanceBench." I want to be exact about what is knowable there. The evaluated subset is not disclosed. Neither is the question count, the judge, nor the generator model. FinanceBench's full open-source set and its 150-question human-evaluated sample are different things, and a figure quoted without saying which it refers to cannot be compared against the paper's numbers. The harder problem is arithmetic. The oracle ceiling from Step 2 is 85%. > Retrieval cannot lift a system above the accuracy its generator reaches when it is handed the right pages. So a system reporting 98.7% is reporting something that cannot be attributed to retrieval alone. Either the generator is substantially better than the GPT-4-Turbo the paper evaluated, or the evaluation subset differs, or the judging criteria differ. Probably some of each. The claim may well be honest; it does not support the inference that retrieval architecture produced it. There is a third scoping issue the README itself states. The open-source package uses standard PDF parsing, while enhanced OCR and tree building are offered as a cloud service. The open-source package and the thing that scored 98.7% are not the same artifact. It is also worth checking where the figure does not appear: pageindex.ai carries no benchmark claims, so this number circulates in write-ups rather than sitting on the project's own site as something it stands behind. On cost there is nothing to cite, because no latency, throughput or cost figures are published anywhere I can find. *What I can say comes from the structure rather than from a measurement.* Every query costs multiple sequential LLM calls to walk the tree. Descent cannot be parallelised, because each level depends on the selection made at the level above. Ingestion requires an LLM pass over the whole document. Those are properties of the design, not numbers I have taken. The remaining limitation I've seen repeated, that scaling past a few hundred documents is unproven, I'm passing on secondhand and have not tested at that size. **The tell:** take any RAG accuracy claim and subtract the benchmark's oracle condition. Whatever sits above that line was not produced by retrieval, and if the write-up doesn't say what produced it, you don't know what you're buying. --- ### Step 5 — Make supersession a relation, not a ranking signal Back to the addendum that started this. A hybrid search that surfaces a superseded clause and its addendum-revised replacement with equal confidence and no temporal ordering is worse than returning nothing, because it produces a defensible-looking answer nobody will re-check. Supersession has to be a relation in the index, enforced at query time, rather than a signal the ranker is trusted to pick up from wording. ```sql -- index_schema.sql -- The document tree, with supersession as a first-class relation rather than -- something the ranker is expected to figure out. CREATE TABLE spec_node ( node_id text PRIMARY KEY, project_id text NOT NULL, parent_id text REFERENCES spec_node(node_id), csi_number text, -- '03 11 00' division int GENERATED ALWAYS AS (substring(csi_number,1,2)::int) STORED, section_part text, -- 'Part 2 Products' title text NOT NULL, page_start int NOT NULL, page_end int NOT NULL, source_doc_id text NOT NULL REFERENCES source_doc(doc_id), summary text ); CREATE TABLE source_doc ( doc_id text PRIMARY KEY, project_id text NOT NULL, doc_type text NOT NULL, -- spec | drawing | addendum | change_order -- | agreement | schedule issued_on date NOT NULL, -- among same-level docs, latest controls sheet_no text -- drawings only ); -- The relation that has to be explicit. An addendum does not simply mention a -- section; it replaces part or all of it as of its issue date. CREATE TABLE supersedes ( superseding_node text REFERENCES spec_node(node_id), superseded_node text REFERENCES spec_node(node_id), effective_on date NOT NULL, scope text NOT NULL, -- 'full' | 'partial' PRIMARY KEY (superseding_node, superseded_node) ); CREATE INDEX ON spec_node (project_id, csi_number); CREATE INDEX ON supersedes (superseded_node); ``` The `doc_type` column is doing more work than it looks like, because construction contracts usually specify an order of precedence among documents. A representative ordering puts change orders and written amendments first, then the Agreement, then drawings, where large scale governs over small scale, then specifications and addenda issued before execution, then owner-furnished information, then other listed Contract Documents. Among documents at the same level, the latest dated one controls. Federal contracts have their own rule in FAR 52.236-21: "In the case of difference between drawings and specifications, the specifications shall govern." The nuance that stops this being a clean algorithm is that the AIA generally advises against precedence clauses at all, on the grounds that they remove the architect's interpretive autonomy. Whether a given project even has an order of precedence is project-specific, and a system that hardcodes one is asserting a contract term that may not exist in that contract. *Every specific in those two paragraphs is secondhand, taken from summaries rather than from contract documents or standard forms.* That flag is the argument, not a hedge. If I cannot establish from secondary reading which precedence rule governs a particular project, a retrieval system reading the same material certainly cannot, and it has no business behaving as though it can. **The tell:** ask your index which of two retrieved clauses is currently in force. If the only answer available is a similarity score, the system cannot tell a live requirement from a dead one. --- ### Step 6 — Verify the citation, then hand the conflict to a person ALCE (arXiv 2305.14627) is where the two metrics I now consider mandatory come from, *as they are commonly described rather than as I read them in the paper*. Citation recall asks whether every statement is fully supported by its cited passages. Citation precision asks whether each individual citation actually supports the claim it is attached to. Both are reported to be computed with an NLI model and validated against human judgment, and the definitions are what I use regardless of how the original evaluation was run. ```python # verify.py from dataclasses import dataclass @dataclass class Span: """Raw offsets kept from the parser all the way through. If chunking loses these, verification becomes a fuzzy string match against the source.""" doc_id: str page: int char_start: int char_end: int text: str @dataclass class Claim: text: str cites: list[Span] def verify(claim: Claim, nli) -> dict: """Precision: does each cited span entail the claim on its own or jointly? Recall: is the claim fully supported by the union of its cited spans?""" per_cite = [nli(premise=s.text, hypothesis=claim.text) for s in claim.cites] joint = nli( premise="\n".join(s.text for s in claim.cites), hypothesis=claim.text, ) supporting = sum(1 for r in per_cite if r == "entailment") return { "citation_precision": supporting / max(len(claim.cites), 1), "citation_recall": 1.0 if joint == "entailment" else 0.0, "contradicted_by": [ s for s, r in zip(claim.cites, per_cite) if r == "contradiction" ], } ``` Keeping raw span offsets from the parser end to end is what makes that check possible, and it is what most chunking pipelines quietly destroy. > A citation that exists but does not entail the claim is worse than no citation, because it manufactures verifiability. It invites trust in a link nobody will follow, and it survives every review that consists of checking whether citations are present. The `contradicted_by` field is where this stops being a metric and becomes a control-flow decision. An agent that resolves a conflict between a specification and a drawing on its own is doing something with contractual and financial consequences, potentially change-order-sized. The correct behaviour when it detects a conflict is to surface it, cite both sources with section number and sheet number, state the applicable precedence clause if the contract contains one, and route the question to an RFI. ```mermaid flowchart TD Q[question] --> R[tree descent] R --> V{cites entail?} V -->|no| REF[refuse] V -->|contradiction| RFI[surface both, route to RFI] V -->|yes| S{one doc in force?} S -->|no| RFI S -->|yes| A[answer with citation] ``` The contractor's obligation on discovering a conflict between contract documents is to report it to the architect through an RFI, not to pick the interpretation that seems more sensible. **The tell:** count how many of your system's outputs in a week were conflicts escalated rather than resolved. If that number is zero, either your corpus has no contradictions in it, or your system is silently picking. --- ### What breaks it **The parser is the real ceiling and nobody measures it separately.** OmniDocBench (CVPR 2025) is the resource I'd point at, *described secondhand* as 981 PDF pages across 9 document types, with a three-level evaluation protocol covering end-to-end performance, single-module performance for OCR, layout detection, table recognition, formula recognition and reading order, and attribute-based robustness. Its exact scale is not the point I need it for. The point holds without it: a RAG system's accuracy ceiling is its parser's accuracy. If table recognition drops a row from a door schedule, no amount of retrieval quality recovers it, and the failure gets attributed to the model. **The contradiction is between four modalities, and the literature only handles one.** Contradiction detection is framed as natural language inference, with its three-way entailment, neutral and contradiction judgement. The dataset lineage runs from SNLI, MNLI and ANLI, too shallow for this, through ContractNLI at the clause level, to ContraDoc and ECON at document scale. arXiv 2504.00180 was the most directly applicable, because it separates self-contradictions within one document from pairwise contradictions across documents, and those need different handling. Construction still breaks the framing: the same requirement appears as prose in the specification, as a dimension on a drawing, as a row in a schedule, and as a revision in an addendum. Detecting that a dimension callout disagrees with a spec paragraph is not a text entailment problem in any form the benchmarks cover. **The model's training data quietly overrules the document.** The taxonomy worth carrying into design splits inter-document conflicts, where two retrieved passages contradict each other, from parametric-contextual conflicts, where retrieved evidence contradicts what the model learned during training. The second is the quieter danger in a technical domain. A model that has absorbed a typical concrete compressive strength will smooth over a spec calling for something unusual, and produce an answer that is right about concrete in general and wrong about this building. --- ### When this architecture is the wrong choice Most of the machinery above is bought with sequential LLM calls at query time and an LLM pass per document at ingestion, and there are no published cost figures to size that against. Reach for it only when the conditions hold. - **The whole document fits in context and a human reads every answer.** Long context scored 79% on that benchmark. It also converted refusals into wrong answers, so this only works while a person is checking, and it stops working the moment nobody is. - **The documents have no authored hierarchy.** Email threads, meeting notes, scanned field photos. Tree navigation needs a table of contents to navigate; without one, the tree is induced by an LLM and you are back to approximating structure, at higher cost than embedding it. - **Your errors survive the oracle condition.** Hand the system the exact right pages. If it still gets the answer wrong, that's the 15% the ceiling describes, and it belongs to the generator. Retrieval architecture cannot move it. - **Being wrong is cheap.** If a wrong answer costs a re-ask rather than a rework order, structural citation, supersession tables and entailment checks are overhead. The whole design is priced for the case where somebody acts on the answer without re-reading the source. - **The corpus is larger than a few hundred documents.** I have not tested at that size, and the limitation I've heard repeated is secondhand. That makes it an open question rather than a known failure, which is its own reason to pilot before committing. **The tell is the addendum.** If no document in your corpus can revise another, none of the supersession machinery earns its cost. --- ### The shift I went looking for a better retriever and found that the largest single effect in the numbers I trust was an index boundary, not a model. The second thing I got wrong is that I treated a citation as the end of verification. A citation is a claim about evidence, and the failure that started this was a real section number pointing at a clause that had been revised. Every review that checks whether citations are present passes that answer. - **Split the index before you tune the model.** 19% to 50% on FinanceBench, same model, same questions. - **Subtract the oracle ceiling from every accuracy claim you read.** What is left above the line was not produced by retrieval. - **Make supersession and precedence explicit relations.** A ranker inferring them from wording is guessing about a contract term. The concrete next step: take ten answers your system produced last week, and check each cited span against the claim it is attached to and against the issue date of the document it came from. > A system that silently picks between contradictory documents is not being helpful. It is quietly assuming liability the contract assigns elsewhere. --- Source: https://himanshuat.com/blogs/when-the-documents-disagree ======================================================================== --- title: "80/20 Your Life!" author: "Damon Zahariades" category: "Self-Help" url: "https://himanshuat.com/books/80-20-your-life" --- # 80/20 Your Life! *Damon Zahariades, 2018* I'm reading this one right now, so the notes and the bits worth keeping will land here once I'm done. It's on my [currently-reading shelf](https://www.goodreads.com/review/list/96919145?shelf=currently-reading) in the meantime. --- Source: https://himanshuat.com/books/80-20-your-life ======================================================================== --- title: "A Clockwork Orange" author: "Anthony Burgess" category: "Psychology" url: "https://himanshuat.com/books/a-clockwork-orange" --- # A Clockwork Orange *Anthony Burgess, 1962* ## Summary The first third of the book is meant to repel you. Alex and his "droogs" speak in Nadsat — Burgess's invented Russian-English teenage slang — and commit a string of crimes whose violence is rendered in a sing-song rhythm that makes the prose itself feel implicated. This is the trap. By the time Alex is caught and the state decides to "cure" him via the Ludovico Technique — a conditioning protocol that makes him physically ill at the thought of violence — you are uncomfortably ready to support whatever fixes him. What Burgess does next is the point of the book. The cure works. Alex is rendered incapable of harm. He is also rendered incapable of choice. The novel's central question — voiced by a prison chaplain almost in passing — is whether a man who cannot choose evil is morally good, or merely a clockwork orange: organic on the outside, mechanical inside. The book's answer is unambiguous, and unpopular. The chaplain calls it goodness imposed from without, which is no goodness at all. The American edition for decades omitted the final chapter, in which Alex begins, on his own, to age out of violence. Burgess hated this. The omission turned the book from a study of moral development into a tract about the impossibility of redemption. The full version is more honest and more difficult: it argues that genuine ethical growth is slow, internal, embarrassing, and incompatible with the state's preferred timetable. It is one of those books people remember as being about violence. It is really about whether free will is worth preserving even when its outputs are monstrous, and what kind of society quietly trades the answer away. ## Quotes > "When a man cannot choose he ceases to be a man." > "The not-self cannot have the bad, meaning they of the government and the judges and the schools cannot allow the bad because they cannot allow the self." > "Goodness is something to be chosen. When a man cannot choose he ceases to be a man." > "What's it going to be then, eh?" ## Why it stayed with me It's the rare novel that puts a serious philosophical question — Kant on autonomy, basically — inside a propulsive crime story and makes you feel both at once. The recurring opening line, "what's it going to be then, eh?", asks the moral question without ever naming it. --- Source: https://himanshuat.com/books/a-clockwork-orange ======================================================================== --- title: "Animal Farm" author: "George Orwell" category: "Philosophy" url: "https://himanshuat.com/books/animal-farm" --- # Animal Farm *George Orwell, 1945* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Animal+Farm+Orwell) in the meantime. --- Source: https://himanshuat.com/books/animal-farm ======================================================================== --- title: "Anthem" author: "Ayn Rand" category: "Philosophy" url: "https://himanshuat.com/books/anthem" --- # Anthem *Ayn Rand, 1938* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Anthem+Ayn+Rand) in the meantime. --- Source: https://himanshuat.com/books/anthem ======================================================================== --- title: "Eat, Pray, Love" author: "Elizabeth Gilbert" category: "Self-Help" url: "https://himanshuat.com/books/eat-pray-love" --- # Eat, Pray, Love *Elizabeth Gilbert, 2006* ## Summary After a marriage and an affair both end badly, Gilbert leaves New York for a year — four months in Italy learning to eat without guilt, four months in an Indian ashram learning to sit with her own mind, four months in Bali learning to love without performing. The structure is tidy. The reader, by halfway through, notices it is too tidy, and that the book is honest about being tidy. Gilbert keeps saying so out loud. The self-awareness is the saving grace. The book has been mocked into a kind of pop-culture shorthand for self-indulgent recovery tourism — sometimes fairly. But underneath the marketability there is a more careful project. The Italy section is about whether a woman is allowed to enjoy food alone. The India section is the most demanding stretch, where the prose stops being pretty and starts being honest about how unpleasant it is to sit still with the contents of your own head. The Bali section is the one that ages the worst, but contains the truest sentence in the book: she meets someone, allows it to be ordinary, and stops trying to make it mean more than it does. What works is that Gilbert is willing to be unflattering about herself in a genre that almost never is. She is petty, vain, anxious about how she sounds, and she puts those qualities on the page. The self-help books that age best are the ones where the author is visibly a real person, not a wisdom-vending machine. Gilbert is a real person, and that's most of why this book keeps getting read fifteen years on. I'd recommend it with a caveat: read it for the methodology, not the destination. The structure — give yourself a year, change one thing each season, write down what happened — is more durable than the prose. The prose is from 2006 and reads like it. The framework is older and reads like it could be from anywhere. ## Quotes > "You need to learn how to select your thoughts just the same way you select your clothes every day. This is a power you can cultivate." > "People think a soul mate is your perfect fit, and that's what everyone wants. But a true soul mate is a mirror, the person who shows you everything that is holding you back." > "In the end, I've come to believe in something I call 'The Physics of the Quest.'" > "Ruin is a gift. Ruin is the road to transformation." ## Why it stayed with me The "select your thoughts like clothes" image is hokey on its face and entirely correct in practice. I still use it as a check on whether a bad mood is a circumstance I have to live with or a wardrobe choice I keep making. --- Source: https://himanshuat.com/books/eat-pray-love ======================================================================== --- title: "Essays in Love" author: "Alain de Botton" category: "Psychology" url: "https://himanshuat.com/books/essays-in-love" --- # Essays in Love *Alain de Botton, 1993* ## Summary The conceit is small on purpose: a narrator meets a woman named Chloe on a flight from Paris to London, they fall in love, they have a relationship, it ends. The book is what happens inside that ordinary arc. De Botton interrupts the narrative every few pages to footnote what he just felt — Marxism in matters of romantic worth, the philosophy of "she would never wear that," the precise mechanics of jealousy at a dinner party — and the effect is that an entirely unremarkable affair becomes the most carefully examined relationship in print. What de Botton sees clearly, and most novelists pretend not to, is that love is largely an act of imagination performed on inadequate evidence. We meet someone, decide quickly, and then spend months retrofitting reality to match the decision. He calls it "the Marxism of love" — the suspicion, borrowed from Groucho, that anyone who would have us cannot possibly be worth having. This is funny and also true, and the book's best gift is that it makes you laugh at your own romantic life while taking it more seriously than before. The later chapters turn darker without losing their lightness. The death of love, in his telling, isn't betrayal or boredom — it's the slow, asymmetric fading of the spell, where one person continues to find the other miraculous while the other stops being able to. He is unsentimental about how badly people behave when this happens, including the narrator, including (it is heavily implied) the author himself. It's a book I keep handing to people in their twenties, with the warning that it will not make them better at love but will make them better at noticing what they're doing. That, I think, is most of what philosophy can do for ordinary life. ## Quotes > "The longing for a destiny is nowhere stronger than in our romantic life." > "We fall in love because we long to escape from ourselves with someone as ideal as we are corrupt." > "Every fall into love involves the triumph of hope over self-knowledge." > "Intimacy is a process by which two people gradually come to discover and reveal the fullness of their respective imperfections." ## Why it stayed with me It's the only book I've read that takes the small, embarrassing thoughts of a relationship — the ones you wouldn't say out loud — and treats them as legitimate philosophical material. Reading it feels less like being taught and more like being caught. --- Source: https://himanshuat.com/books/essays-in-love ======================================================================== --- title: "Factfulness" author: "Hans Rosling" category: "Psychology" url: "https://himanshuat.com/books/factfulness" --- # Factfulness *Hans Rosling, 2018* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Factfulness+Rosling) in the meantime. --- Source: https://himanshuat.com/books/factfulness ======================================================================== --- title: "Flowers for Algernon" author: "Daniel Keyes" category: "Psychology" url: "https://himanshuat.com/books/flowers-for-algernon" --- # Flowers for Algernon *Daniel Keyes, 1966* ## Summary The book is told entirely as Charlie Gordon's progress reports. Charlie starts the novel with an IQ of 68 and ends it back where he started, having spent the intervening months as one of the most brilliant minds on the planet. The prose changes with him — misspelled and warmly literal at the opening, dense and self-aware at the peak, then breaking down again as the procedure fails. The form is the argument. What Keyes gets right is that intelligence is not a neutral upgrade. As Charlie's IQ climbs he doesn't become happier; he becomes lonelier, more accurate about the people who used to love him, less able to tolerate the kindness he used to mistake for friendship. The same mother who couldn't accept his disability cannot accept his intelligence either. Algernon, the lab mouse who got the procedure first, is the foreshadow Charlie can't bring himself to read until it's too late. It would be easy to read the book as a tragedy about a man losing what he briefly had, but Keyes is after something more uncomfortable. The version of Charlie who knows the most is not the version most able to live. Knowledge, in this book, is not the same as wisdom and not the same as peace. The reader is left with a question that good psychology asks and rarely answers: what would you actually do with a doubled mind, and would you still recognize yourself at the end of it? I read this young and remembered it as sad. I reread it as an adult and remembered it as a study of how much of our identity is held up by the people around us, and how quickly that scaffolding shows when the self underneath starts to change. ## Quotes > "I don't know what's worse: to not know what you are and be happy, or to become what you've always wanted to be, and feel alone." > "Intelligence and education that hasn't been tempered by human affection isn't worth a damn." > "Now I understand one of the important reasons for going to college and getting an education is to learn that the things you've believed in all your life aren't true, and that nothing is what it appears to be." > "Even a feeble-minded man wants to be like other men." ## Why it stayed with me It's the rare book where the prose teaching the character's interior is also teaching yours. By the time Charlie's spelling begins to break down at the end, you don't need to be told what is happening — you feel it on the page. --- Source: https://himanshuat.com/books/flowers-for-algernon ======================================================================== --- title: "How Doctors Think" author: "Jerome Groopman" category: "Psychology" url: "https://himanshuat.com/books/how-doctors-think" --- # How Doctors Think *Jerome Groopman, 2007* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=How+Doctors+Think+Groopman) in the meantime. --- Source: https://himanshuat.com/books/how-doctors-think ======================================================================== --- title: "Hunger" author: "Knut Hamsun" category: "Psychology" url: "https://himanshuat.com/books/hunger" --- # Hunger *Knut Hamsun, 1890* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Hunger+Hamsun) in the meantime. --- Source: https://himanshuat.com/books/hunger ======================================================================== --- title: "Invisible Man" author: "Ralph Ellison" category: "Psychology" url: "https://himanshuat.com/books/invisible-man" --- # Invisible Man *Ralph Ellison, 1952* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Invisible+Man+Ellison) in the meantime. --- Source: https://himanshuat.com/books/invisible-man ======================================================================== --- title: "The Strange Case of Dr. Jekyll and Mr. Hyde" author: "Robert Louis Stevenson" category: "Psychology" url: "https://himanshuat.com/books/jekyll-and-hyde" --- # The Strange Case of Dr. Jekyll and Mr. Hyde *Robert Louis Stevenson, 1886* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Jekyll+and+Hyde) in the meantime. --- Source: https://himanshuat.com/books/jekyll-and-hyde ======================================================================== --- title: "Life is Elsewhere" author: "Milan Kundera" category: "Psychology" url: "https://himanshuat.com/books/life-is-elsewhere" --- # Life is Elsewhere *Milan Kundera, 1969* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Life+is+Elsewhere+Kundera) in the meantime. --- Source: https://himanshuat.com/books/life-is-elsewhere ======================================================================== --- title: "Man's Search for Meaning" author: "Viktor Frankl" category: "Psychology" url: "https://himanshuat.com/books/mans-search-for-meaning" --- # Man's Search for Meaning *Viktor Frankl, 1946* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Mans+Search+for+Meaning+Frankl) in the meantime. --- Source: https://himanshuat.com/books/mans-search-for-meaning ======================================================================== --- title: "Meditations" author: "Marcus Aurelius" category: "Philosophy" url: "https://himanshuat.com/books/meditations" --- # Meditations *Marcus Aurelius, 180* ## Summary *Meditations* was never meant to be read. Marcus Aurelius wrote it for himself — twelve short books of reminders, addressed almost entirely in the second person, hammering on the same handful of ideas until they stuck. That accidental quality is what makes it the most durable piece of philosophy I've encountered. There is no thesis to defend, no audience to impress; just a man at the top of the world telling himself, again and again, that the world's opinion of him does not matter and that he will be dead soon and that the only thing he controls is how he chooses to act in the next hour. The philosophy is late Stoicism, but lighter than that label suggests. The recurring moves are: separate what you can control from what you cannot; notice that almost everything is in the second category; treat other people's failings as the inevitable output of their training rather than personal injuries; remember, constantly, that you will die. None of this is news, then or now. What makes it land is the voice — patient, slightly tired, occasionally angry with itself, never theatrical. You are reading a man trying to manage himself, not a philosopher selling a system. Modernity does a lot to this book that it shouldn't. The "stoicism" tradition has been flattened into productivity advice and cold-shower discourse, which Marcus would have found embarrassing. The real argument is closer to a kind of cosmic humility: you are a small thing in a large universe, almost everything you worry about will be forgotten within a generation, and the appropriate response to this is not despair but a focused, ethical attention to the small piece of life you actually have. I reread the first book — his catalogue of what he learned from each person who raised him — about once a year. It is the most generous thing in the text, and it teaches the central lesson sideways: that a life is measured mostly by what you absorbed from the people who took the trouble to be near you. ## Quotes > "You have power over your mind — not outside events. Realize this, and you will find strength." > "When you arise in the morning, think of what a precious privilege it is to be alive — to breathe, to think, to enjoy, to love." > "The happiness of your life depends upon the quality of your thoughts." > "Waste no more time arguing about what a good man should be. Be one." ## Why it stayed with me It's the rare ancient book that gets shorter, not longer, with each reread — most of the lines you remembered the first time were already the right ones. The rest is decoration. --- Source: https://himanshuat.com/books/meditations ======================================================================== --- title: "Moonwalking with Einstein" author: "Joshua Foer" category: "Self-Help" url: "https://himanshuat.com/books/moonwalking-with-einstein" --- # Moonwalking with Einstein *Joshua Foer, 2011* ## Summary Foer set out to write a magazine piece about memory athletes and ended up training as one. The book is the story of that year — the people he met, the techniques he learned, and the slowly dawning realization that the people who can recite a deck of cards in under two minutes are not, by any meaningful measure, smarter than the rest of us. They've just practiced. The central technique is ancient: the "memory palace," a method described by Cicero and the anonymous author of *Rhetorica ad Herennium* almost two millennia ago. You convert what you want to remember into vivid, often absurd images, and you place those images along a familiar mental route — your childhood home, your walk to school. Recall becomes a stroll. Foer makes the case that this isn't a parlor trick but a lost civic skill, something Western education quietly stopped teaching once books got cheap and Google got fast. The more interesting argument is what we lost when we offloaded memory to paper and silicon. Foer sits with Kim Peek, the inspiration for *Rain Man*; he interviews S, the Russian journalist whose memory was so total he couldn't forget; he profiles EP, a man whose hippocampus has been destroyed and who lives in a permanent thirty-second present. Together these portraits sketch the strange shape of a faculty most of us never think about until it begins to fail. It's a self-help book disguised as a memoir disguised as journalism, and the disguise works. You finish it convinced that "I have a bad memory" is roughly as honest as "I'm bad at running" — true, probably, but only because you haven't trained. ## Quotes > "How we perceive the world and how we act in it are products of how and what we remember. We're all just a bundle of habits shaped by our memories. And to the extent that we control our lives, we do so by gradually altering those habits." > "The brain is a costly organ. Though it accounts for only 2 percent of the body's mass, it uses up a fifth of all the oxygen we breathe... It's there for a reason. And that reason is to remember." > "Monotony collapses time; novelty unfolds it. You can exercise daily and eat healthily and live a long life, while experiencing a short one." > "The more we remember, the better we are at processing the world. And the better we are at processing the world, the more we can remember about it." ## Why it stayed with me The line about monotony and novelty changed how I think about whether a year was a "good" one. And the memory palace stuff actually works — I still use it for grocery lists and the occasional speech. --- Source: https://himanshuat.com/books/moonwalking-with-einstein ======================================================================== --- title: "No Exit and Three Other Plays" author: "Jean-Paul Sartre" category: "Philosophy" url: "https://himanshuat.com/books/no-exit" --- # No Exit and Three Other Plays *Jean-Paul Sartre, 1944* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=No+Exit+Sartre) in the meantime. --- Source: https://himanshuat.com/books/no-exit ======================================================================== --- title: "On the Shortness of Life" author: "Seneca" category: "Philosophy" url: "https://himanshuat.com/books/on-the-shortness-of-life" --- # On the Shortness of Life *Seneca, 49* ## Summary Seneca's essay opens with one of the cleanest reversals in philosophy. People complain that life is too short, he says. They are wrong. Life is long enough — they have just wasted most of it on other people's business, on flattery, on the careful management of a reputation no one but them is keeping track of. "It is not that we have a short time to live," runs the famous line, "but that we waste a lot of it." Once you accept the premise, the rest of the essay is just Seneca methodically listing the ways we do this. The book is structured as a letter to his father-in-law, Paulinus, who managed Rome's grain supply — a real job with real consequences, and Seneca is gently suggesting Paulinus quit it. Not because the work is unimportant, but because no amount of public usefulness compensates for a life in which the man himself is never at home. The argument generalizes: even meaningful busyness, if it never lets up, is a form of theft. What's surprising on a careful read is how much of the essay is about *attention* rather than *time*. The truly long-lived person, in Seneca's framing, is not the one who lives many years but the one who is present for the years he has — who keeps a relationship with his own past, his own thinking, his own friends. The "preoccupied" man, by contrast, has no past he can return to without flinching and no future he isn't busy outrunning. This is essentially the modern observation about phones and attention, written down before Pompeii fell. It's also short. You can read it in an hour. That's not incidental — the form is the argument. A book that took a week to read would have undermined itself. ## Quotes > "It is not that we have a short time to live, but that we waste a lot of it. Life is long enough, and a sufficiently generous amount has been given to us... if it were all well invested." > "You are living as if destined to live forever; your own frailty never occurs to you; you don't notice how much time has already passed." > "Everyone hurries his life on and suffers from a yearning for the future and a weariness of the present." > "Of all people only those are at leisure who make time for philosophy, only those are really alive. For they not only keep a good watch over their own lifetimes, but they annex every age to theirs." ## Why it stayed with me The line about being "preoccupied" describes about ninety percent of how I've spent my own attention, and I haven't found a better diagnosis since. The essay is a corrective you can return to whenever you notice the symptom. --- Source: https://himanshuat.com/books/on-the-shortness-of-life ======================================================================== --- title: "Prometheus Rising" author: "Robert Anton Wilson" category: "Psychology" url: "https://himanshuat.com/books/prometheus-rising" --- # Prometheus Rising *Robert Anton Wilson, 1983* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Prometheus+Rising+Wilson) in the meantime. --- Source: https://himanshuat.com/books/prometheus-rising ======================================================================== --- title: "The Age of Reason" author: "Jean-Paul Sartre" category: "Philosophy" url: "https://himanshuat.com/books/the-age-of-reason" --- # The Age of Reason *Jean-Paul Sartre, 1945* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Age+of+Reason+Sartre) in the meantime. --- Source: https://himanshuat.com/books/the-age-of-reason ======================================================================== --- title: "The Analects" author: "Confucius" category: "Philosophy" url: "https://himanshuat.com/books/the-analects" --- # The Analects *Confucius, -500* ## Summary The Analects is not a treatise. It is a notebook of remembered moments — Confucius answering a question, refusing to answer a question, sighing about a former student, correcting a duke. The text has no plot and no system. What it has is a slowly accumulating portrait of a person who took ethics personally enough that the people around him kept writing down what he said. The central idea, untranslatable but circled by every chapter, is *ren* — usually rendered as benevolence or humaneness. It is not a feeling. It is a practiced disposition, built daily, in small acts, in the presence of family and strangers and superiors and the dead. The closest Western analogue is virtue ethics, but Confucius is more specific: the work of becoming good is the work of becoming a particular kind of son, a particular kind of friend, a particular kind of person who can be relied on. Abstractions are not interesting to him. Particulars are. Modernity reads this book at its peril. The instinct is to extract the lines that double as fortune cookies and discard the rest. But the discarded rest is the actual book — the rituals, the music, the proper way to bow, the seasonal sacrifices. The argument is that you do not become wise by knowing wise sentences. You become wise by behaving carefully, over decades, in tens of thousands of small encounters. The reason the text feels demanding rather than soothing is that it does not let you take the shortcut you came for. I keep a copy on the desk and open it at random when I notice myself getting clever. It is a corrective book — the kind you finish less convinced of yourself than when you started. That, I have come to think, is the actual point. ## Quotes > "When you know a thing, to hold that you know it; and when you do not know a thing, to allow that you do not know it — this is knowledge." > "The superior man may indeed have to endure want, but the mean man, when he is in want, gives way to unbridled license." > "Have no friends not equal to yourself." > "It does not matter how slowly you go as long as you do not stop." ## Why it stayed with me The line about admitting what you do not know is, I think, the single most important sentence anyone has ever written about how to think. I run it as a private check on my own opinions, and lose more arguments because of it than I would otherwise. --- Source: https://himanshuat.com/books/the-analects ======================================================================== --- title: "The Body Keeps the Score" author: "Bessel van der Kolk" category: "Psychology" url: "https://himanshuat.com/books/the-body-keeps-the-score" --- # The Body Keeps the Score *Bessel van der Kolk, 2014* ## Summary Van der Kolk's central claim is unfashionably literal: trauma is not just a story you tell yourself badly, it is a thing that lives in the body. Decades of work with combat veterans, abuse survivors, and children in chronic stress led him to a picture of the traumatized nervous system as a smoke alarm wired too sensitively — flipping between hyperarousal and shutdown long after the original danger has passed. The book is part history of psychiatry, part neuroscience tour, part case file. What makes it more than a clinical text is the way he refuses to settle for any one explanation. He walks through the rise and limits of pharmacology, the politics that kept PTSD out of the diagnostic manual until 1980, and the quiet evidence behind treatments the mainstream still treats as fringe — EMDR, yoga, theater, neurofeedback, internal family systems. The thread that holds it together is a simple, hard-won idea: healing requires re-inhabiting the body, not just rethinking the past. The case studies do most of the moral work. A Vietnam veteran who can't read his children a bedtime story without dissociating. A young woman whose body remembers what her mind has carefully forgotten. They aren't there for shock value — they're there to make the neuroscience refuse to stay abstract. By the time he gets to the chapters on treatment, the reader has already understood, viscerally, why talking alone isn't enough. I came in skeptical of the genre and left convinced that this is one of the few books on the subject that earns its breadth. It changed how I read other people, and probably how I read myself. ## Quotes > "As long as you keep secrets and suppress information, you are fundamentally at war with yourself... The critical issue is allowing yourself to know what you know." > "Being able to feel safe with other people is probably the single most important aspect of mental health; safe connections are fundamental to meaningful and satisfying lives." > "Neuroscience research shows that the only way we can change the way we feel is by becoming aware of our inner experience and learning to befriend what is going on inside ourselves." > "Trauma is not the story of something that happened back then. It's the current imprint of that pain, horror, and fear living inside people." ## Why it stayed with me It reframes "mental" health as something the body has been trying to tell you about for a long time. After this book it's harder to take seriously any therapy framework that treats the neck-down as an inconvenience. --- Source: https://himanshuat.com/books/the-body-keeps-the-score ======================================================================== --- title: "The Brothers Karamazov" author: "Fyodor Dostoevsky" category: "Psychology" url: "https://himanshuat.com/books/the-brothers-karamazov" --- # The Brothers Karamazov *Fyodor Dostoevsky, 1880* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Brothers+Karamazov) in the meantime. --- Source: https://himanshuat.com/books/the-brothers-karamazov ======================================================================== --- title: "The Celestine Prophecy" author: "James Redfield" category: "Self-Help" url: "https://himanshuat.com/books/the-celestine-prophecy" --- # The Celestine Prophecy *James Redfield, 1993* ## Summary Redfield's book is technically a novel, but only barely. A first-person narrator chases a mysterious manuscript through the Peruvian rainforest, and at each stop along the way someone explains another of the manuscript's nine "Insights" to him. The plot is mostly scaffolding. What people actually came for — and what made the book a multi-year fixture on the bestseller list — is the framework: a vaguely New Age, vaguely Jungian account of how human beings exchange psychic energy and why coincidence might be worth a second look. I will say upfront that the prose is not why anyone keeps reading. The Insights are. They argue, in order, that coincidences are signals, that history is moving toward a more spiritual culture, that humans operate inside a kind of energy field, that most conflict is unconscious competition for that energy, and that "control dramas" — intimidator, interrogator, aloof, poor-me — are the four scripts most of us run on other people without noticing. The last part is the actually useful one. Once you've named the four dramas, you start to see them in your own family within a week. What the book gets right, even when the metaphysics overreaches, is the diagnosis that modern life has flattened a class of experience — meaningful coincidence, intuition, presence — that used to be the province of religion and is now treated as embarrassing. Redfield's solution is not particularly rigorous, but the question is. Whether or not you buy the energy-field language, the underlying suggestion — that the universe is more alive than your default mode assumes, and that paying attention to it changes you — has held up better than the prose has. I read this young, dismissed it, and then noticed years later that the "control dramas" chapter was still rattling around in my head. That's usually the test of whether a book did something to you, regardless of whether you'd recommend it. ## Quotes > "Where attention goes, energy flows." > "Once we question the higher meaning of life, the next step is to be vigilant — to be on the lookout for the actual experience of new things... When this awareness occurs, then we are connecting with the source of energy in the universe." > "Every event has meaning. There are no chance events." > "The truth is, we cannot evolve any faster than we can release our control dramas." ## Why it stayed with me The "control dramas" framework is one of those small naming gifts that outlives the book it came from. I don't endorse the cosmology, but I've used the vocabulary in more honest conversations than I can count. --- Source: https://himanshuat.com/books/the-celestine-prophecy ======================================================================== --- title: "The Death of Ivan Ilyich" author: "Leo Tolstoy" category: "Psychology" url: "https://himanshuat.com/books/the-death-of-ivan-ilych" --- # The Death of Ivan Ilyich *Leo Tolstoy, 1886* ## Summary The novella's first chapter shows Ivan Ilyich's funeral; the rest of the book walks back through the life that produced the body in the coffin. Tolstoy does this on purpose. We know the ending. The question is what was inside it. What was inside it, in Tolstoy's account, is the most ordinary kind of self-deception. Ivan rises through the courts, marries the appropriate woman, decorates the appropriate apartment, and at no point in any of this is he living a life he would have chosen had he been paying attention. The terrifying thing is how socially competent the deception is — his colleagues recognize it, his wife recognizes it, and they all collaborate to keep him from noticing. When the illness arrives, the conspiracy of "you'll be fine, dear" continues right up to the last days. The only person in the book who treats him honestly is a peasant servant named Gerasim, who simply lets the dying man rest his legs on his shoulders and does not pretend the rest. The psychological move at the center of the book is small and devastating. Lying in pain, Ivan asks himself, *what if my whole life was wrong?* — and instead of dismissing the question, he sits with it. The honesty kills him faster, in one sense, and saves him, in another. The last pages are unlike anything else in 19th-century fiction. They are the work of a man who had clearly been thinking about his own death for a long time. I read this expecting a Victorian morality play and got something much closer to a clinical case file. It belongs on a shelf with *The Body Keeps the Score* more than with *Anna Karenina* — a study, not a story. ## Quotes > "Ivan Ilyich's life had been most simple and most ordinary and therefore most terrible." > "What if my whole life has been wrong?" > "It is impossible, but it is." > "He sought his former accustomed fear of death and did not find it. Where was it? What death? There was no fear because there was no death. In place of death there was light." ## Why it stayed with me Most fiction tells you that an examined life is a happier one. Tolstoy is more honest — an examined life is sometimes only happier in its last hour, and the previous decades remain what they were. The novella is what you reach for when the self-improvement bookshelf starts to feel cowardly. --- Source: https://himanshuat.com/books/the-death-of-ivan-ilych ======================================================================== --- title: "The Handmaid's Tale" author: "Margaret Atwood" category: "Philosophy" url: "https://himanshuat.com/books/the-handmaids-tale" --- # The Handmaid's Tale *Margaret Atwood, 1985* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Handmaids+Tale+Atwood) in the meantime. --- Source: https://himanshuat.com/books/the-handmaids-tale ======================================================================== --- title: "The Magic Mountain" author: "Thomas Mann" category: "Philosophy" url: "https://himanshuat.com/books/the-magic-mountain" --- # The Magic Mountain *Thomas Mann, 1924* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Magic+Mountain+Mann) in the meantime. --- Source: https://himanshuat.com/books/the-magic-mountain ======================================================================== --- title: "The Republic" author: "Plato" category: "Philosophy" url: "https://himanshuat.com/books/the-republic" --- # The Republic *Plato, -375* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=The+Republic+Plato) in the meantime. --- Source: https://himanshuat.com/books/the-republic ======================================================================== --- title: "The Screwtape Letters" author: "C.S. Lewis" category: "Philosophy" url: "https://himanshuat.com/books/the-screwtape-letters" --- # The Screwtape Letters *C.S. Lewis, 1942* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Screwtape+Letters) in the meantime. --- Source: https://himanshuat.com/books/the-screwtape-letters ======================================================================== --- title: "The Sense of an Ending" author: "Julian Barnes" category: "Psychology" url: "https://himanshuat.com/books/the-sense-of-an-ending" --- # The Sense of an Ending *Julian Barnes, 2011* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Sense+of+an+Ending+Barnes) in the meantime. --- Source: https://himanshuat.com/books/the-sense-of-an-ending ======================================================================== --- title: "The Stranger" author: "Albert Camus" category: "Philosophy" url: "https://himanshuat.com/books/the-stranger" --- # The Stranger *Albert Camus, 1942* ## Summary The famous opening line — "Mother died today. Or maybe yesterday, I can't be sure." — sets up everything that follows. Meursault is not a monster; he is simply unwilling to pretend to feel things he does not feel. He sees his mother's funeral as hot and inconvenient. He sees his girlfriend's love as pleasant. He sees the killing on the beach as something the sun made him do. The court system does not understand any of this. The trial spends more time on his behavior at the funeral than on the killing itself, which is, in a quiet way, the entire argument of the book. What Camus is doing is making absurdism legible through a man so honest that he refuses to perform the small social emotions everyone else takes for granted. The cost is that society treats him as inhuman. The reward is that he becomes, by the final pages, one of the few characters in 20th-century literature to fully accept the conditions of his existence — including his death — without flinching, without lying, without reaching for a god he doesn't believe in. The book is short and reads short. It is also philosophically heavier than most thousand-page novels. The mechanism is that you spend the first half feeling vaguely uncomfortable with Meursault, and then in the second half the discomfort flips: you realize he is the only character speaking truthfully, and the discomfort was never about him. It was about how much performance the rest of us mistake for sincerity. I came back to it after reading some clinical psychology and found that the diagnostic urge — "what's wrong with this man?" — is exactly what the book is built to dismantle. Camus's answer is that what is "wrong" with Meursault is that he refuses the costume the rest of us are wearing without noticing. ## Quotes > "Mother died today. Or maybe yesterday, I can't be sure." > "I opened myself to the gentle indifference of the world." > "Since we're all going to die, it's obvious that when and how don't matter." > "I had been right, I was still right, I was always right. I had lived my life one way and I could just as well have lived it another. I had done this and I hadn't done that. I hadn't done this thing but I had done another. And so?" ## Why it stayed with me The closing pages are the cleanest defense of living without consolation I've ever read. The whole book builds toward them; it would be a different — lesser — novel without the prison-cell monologue at the end. --- Source: https://himanshuat.com/books/the-stranger ======================================================================== --- title: "The Trial" author: "Franz Kafka" category: "Psychology" url: "https://himanshuat.com/books/the-trial" --- # The Trial *Franz Kafka, 1925* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=The+Trial+Kafka) in the meantime. --- Source: https://himanshuat.com/books/the-trial ======================================================================== --- title: "We" author: "Yevgeny Zamyatin" category: "Philosophy" url: "https://himanshuat.com/books/we" --- # We *Yevgeny Zamyatin, 1924* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=We+Zamyatin) in the meantime. --- Source: https://himanshuat.com/books/we ======================================================================== --- title: "Wind, Sand and Stars" author: "Antoine de Saint-Exupéry" category: "Philosophy" url: "https://himanshuat.com/books/wind-sand-and-stars" --- # Wind, Sand and Stars *Antoine de Saint-Exupéry, 1939* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Wind+Sand+and+Stars) in the meantime. --- Source: https://himanshuat.com/books/wind-sand-and-stars ======================================================================== --- title: "Zorba the Greek" author: "Nikos Kazantzakis" category: "Philosophy" url: "https://himanshuat.com/books/zorba-the-greek" --- # Zorba the Greek *Nikos Kazantzakis, 1946* I've read this one but haven't written up notes for it yet. It's logged on my [Goodreads shelf](https://www.goodreads.com/search?q=Zorba+the+Greek) in the meantime. --- Source: https://himanshuat.com/books/zorba-the-greek ======================================================================== --- title: "Software Development Engineer (SDE) at Tejas AI (YC W25)" company: "Tejas AI (YC W25)" period: "Mar 2026 – Aug 2026" location: "SF, California, USA" url: "https://himanshuat.com/experience/tejas-ai-yc-w25-software-development-engineer-sde" --- # Software Development Engineer (SDE) — Tejas AI (YC W25) *Mar 2026 – Aug 2026 · SF, California, USA* ## About the company Tejas AI builds Elip, an AI career coach that lives where job seekers already are: WhatsApp, with a web app, a Chrome extension and a mobile app around it. The pitch is that looking for work is mostly unglamorous logistics, and most of that can be handled by an agent that remembers you between conversations. ## What I did - Built Elip, an AI career coach on WhatsApp (with web, a Chrome extension, and a mobile app) that helps job seekers fix their resume, build a search profile, and get matched to jobs. - The core was a FastAPI WhatsApp agent, a router model with specialist sub-agents, that parsed resumes and LinkedIn profiles, ran resume reviews, and returned ranked jobs, with cross-session memory across Redis, Postgres, and S3. - Built the search layer over OpenSearch (hybrid keyword and vector search with Bedrock embeddings) and a fleet of Go services that index jobs from around 60 ATS and job boards. - Built the Go backend behind the web and mobile apps, plus a Chrome extension that autofills multi-step job applications and tailors a resume to each posting. ## The longer version The interesting constraint was the surface. WhatsApp is a text box. There is no sidebar to hide settings in, no place to show a form, and no way to make someone read documentation first. Everything the product did had to survive being asked for in one message, by someone who is tired and job hunting. The core was a FastAPI agent shaped as a router with specialist sub-agents behind it. One read resumes and LinkedIn profiles, one ran the review, one returned ranked jobs. Routing sat in code rather than in the model's discretion, because a router that occasionally decides to freestyle is a router that occasionally loses someone's resume. Memory was the part I underestimated. A career conversation is not one session. It is a person coming back three weeks later expecting you to know what they already told you, so state lived across Redis, Postgres and S3 rather than in a context window. The other half was retrieval. I built the search layer on OpenSearch, hybrid keyword and vector with Bedrock embeddings, over a fleet of Go services indexing jobs from around sixty ATS platforms and job boards. Hybrid rather than pure vector for the reason I keep running into: a hard requirement is a filter, not a similarity score. I left in August 2026. Five months, and the thing I took with me is that the hard problems were never the model ones — they were memory, routing determinism, and the difference between a query and a requirement. ## Stack TypeScript, Go, Node.js, Nest.js, PostgreSQL, Redis, OpenSearch, Mem0, LLM, Docker, AWS --- Source: https://himanshuat.com/experience/tejas-ai-yc-w25-software-development-engineer-sde ======================================================================== --- title: "AI Researcher & CTO at AuraX" company: "AuraX" period: "Dec 2024 – Dec 2025" location: "Hyderabad, India" url: "https://himanshuat.com/experience/aurax-ai-researcher-cto" --- # AI Researcher & CTO — AuraX *Dec 2024 – Dec 2025 · Hyderabad, India* ## About the company AuraX was a fashion-tech company incubated at IIIT-Hyderabad's CIE. Fashion e-commerce needs on-model photography for every garment, in every size, on every body type, and a photoshoot per SKU does not scale. We generated that imagery with diffusion models instead. ## What I did - Co-founded AuraX, a fashion-tech startup funded by IIIT-Hyderabad's CIE incubator. - Led research on diffusion models for virtual try-on (VTO), achieving quality that outperformed Google, Stable Diffusion, and leading Indian startups in benchmark comparisons. - Signed commercial deals with Aditya Birla Fashion Group and multiple brands under Mensa Brands. - Built the backend and deployment infrastructure on AWS and ShaktiCloud, and tuned inference latency for production. - Wound the company down as bigger players entered the space and our runway ran short. ## The longer version I co-founded it and ran the research. The problem sounds like image generation and is not. A customer already knows what the shirt looks like, so the garment is a constraint rather than a suggestion, and getting it approximately right is the same as getting it wrong. The gap nobody was serving was non-Western clothing. Open models are trained mostly on Western casual wear, so they handle a t-shirt well and fall apart on a saree, because pleats and draping physics are not something you can approximate from a prior built on tubes of fabric. We built for that case specifically, and it turned out to be the most defensible thing we had. Alongside the research I built the production side: backend and GPU inference across AWS, ShaktiCloud and Modal, with latency tuned until brands would actually put it in front of customers. Modal covered the serverless side, where inference is bursty and paying for an idle GPU makes no sense, while the sustained training and R&D work sat on dedicated nodes. Research quality and shipped quality are different problems, and the second one is less fun and takes longer. It worked. We beat the baselines we tested against, and we signed Aditya Birla Fashion Group and several Mensa Brands labels. Real contracts, real revenue. Then better funded companies entered the space and our runway ran out. I made the call to close it. The lesson I actually took is not about diffusion models: being right about the technology and being right about the business are separate bets, and I had been treating a win in the first as evidence about the second. ## Stack Python, PyTorch, diffusers, LoRA, CUDA, LLM, FastAPI, PostgreSQL, Redis, Docker, AWS, ShaktiCloud, Next.js, Nest.js Company site: https://aurax.co.in --- Source: https://himanshuat.com/experience/aurax-ai-researcher-cto ======================================================================== --- title: "Founder at SynthAiLabs (formerly OpenEdu)" company: "SynthAiLabs (formerly OpenEdu)" period: "Dec 2023 – Feb 2025" location: "Remote" url: "https://himanshuat.com/experience/synthailabs-formerly-openedu-founder" --- # Founder — SynthAiLabs (formerly OpenEdu) *Dec 2023 – Feb 2025 · Remote* ## About the company SynthAiLabs, which started as OpenEdu, was an attempt to compress how long it takes a student to become employable. Part edtech product, part community, aimed at Indian engineering students who had the ability but not the ramp. ## What I did - Researched Mixture-of-Experts (MoE) LLM architectures. - Built learncode (interactive AI lessons) and eduAI (an agent that answers questions about course material). - Grew techos.synthailabs to 11,000 registered students across India. - Co-founded the OpenCodeDevelopers (OCD) society, which grew by word of mouth and helped 800+ students cut their CP/DSA ramp from six months to two. ## The longer version I started it in my second year. The bet was that most of the time students lose is not to difficulty, it is to not knowing what to do next, and that a tight enough structure with people around you closes that gap faster than better material does. On the product side I built learncode, which taught through interactive AI lessons, and eduAI, an agent that answered questions about the actual course material rather than in general. I also spent real time on Mixture-of-Experts architectures, which is where my interest in specialised models rather than one large one comes from. The community part worked better than the product part. techos.synthailabs reached 11,000 registered students across India, and the OpenCodeDevelopers society grew almost entirely by word of mouth, helping more than 800 students cut a six month competitive programming ramp down to about two. Funding was always the constraint and I eventually stepped away. The validation I care about came afterwards: juniors restarted it and kept it running without me. Something that survives its founder leaving was worth more than anything I could have put on a slide. ## Stack TypeScript, Next.js, Node.js, Express, MongoDB, PostgreSQL, Redis, Docker, CI/CD, React, TailwindCSS Company site: https://synthailabs.com --- Source: https://himanshuat.com/experience/synthailabs-formerly-openedu-founder ======================================================================== --- title: "AI Dev at Tublian" company: "Tublian" period: "July 2024 – Dec 2024" location: "Columbus, Ohio" url: "https://himanshuat.com/experience/tublian-ai-dev" --- # AI Dev — Tublian *July 2024 – Dec 2024 · Columbus, Ohio* ## About the company Tublian is a developer platform, with TublianOS as the product and an open-source collaboration programme around it. I came back to it in a second, more engineering-focused stint. ## What I did - Set up Docker-based build and deployment workflows. - Added code obfuscation to protect the source before public deployment. - Contributed to TublianOS, working on performance and the user-facing experience. - Completed a system-design masterclass that used Google's search architecture as its worked example. ## The longer version This was infrastructure work rather than product work. I set up the Docker-based build and deployment workflows, and added code obfuscation so the source was protected before public deployment. The rest was TublianOS itself, working on performance and the parts of the experience users actually touch. The most useful thing I took away was not shipped code. It was a system design masterclass that used Google's search architecture as its worked example, which is the first time I properly understood what a system looks like when it is designed for a scale you cannot hold in your head. ## Stack Python, LLM, Ollama, FastAPI, Docker Company site: https://tublian.com --- Source: https://himanshuat.com/experience/tublian-ai-dev ======================================================================== --- title: "QM Researcher at Stealth" company: "Stealth" period: "March 2024 – June 2024" location: "Kyoto, Japan" url: "https://himanshuat.com/experience/stealth-qm-researcher" --- # QM Researcher — Stealth *March 2024 – June 2024 · Kyoto, Japan* ## About the company A stealth research group in Kyoto working at the intersection of quantum computing and machine learning. The specifics are under NDA, so this entry stays deliberately thin. ## What I did - [NDA], tested ML model deployment on quantum computers using qiskit ## The longer version What I can say: I worked on deploying machine learning models onto quantum computers using Qiskit, and spent most of my time on the gap between what the papers describe and what the hardware will currently tolerate. It is the shortest thing on here and the one that changed my sense of scale the most. Working on something where the tooling is still being invented, in a field where the honest answer to most questions is that nobody knows yet, is a different discipline from shipping software. ## Stack Python, Qiskit, OpenQASM, C/C++ --- Source: https://himanshuat.com/experience/stealth-qm-researcher ======================================================================== --- title: "Open Source Collaborator | Project Mentor at Tublian" company: "Tublian" period: "Dec. 2023 – April 2024" location: "Columbus, Ohio" url: "https://himanshuat.com/experience/tublian-open-source-collaborator-project-mentor" --- # Open Source Collaborator | Project Mentor — Tublian *Dec. 2023 – April 2024 · Columbus, Ohio* ## About the company Tublian's open-source programme pairs developers with real projects and with each other. I joined as a contributor and ended up mentoring. ## What I did - Contributed to several open-source projects and mentored 20+ newer developers. - Helped build DevDocGenie, a RAG-based chatbot for searching documentation. - Reverse-engineered features like Samsung's live call translation as build exercises. ## The longer version I contributed across several open-source projects and mentored more than twenty newer developers. Mentoring turned out to be the part that taught me most, because explaining a decision to someone who will actually go and implement it is a much harder test of whether you understood it than writing it yourself. The build I remember is DevDocGenie, a RAG chatbot for searching documentation, well before retrieval systems were a thing everyone had an opinion about. We also reverse-engineered shipped features as exercises, Samsung's live call translation among them. Taking apart something that already works is an underrated way to learn what the hard parts actually were. ## Stack TypeScript, React, Node.js, GitHub Actions Company site: https://tublian.com --- Source: https://himanshuat.com/experience/tublian-open-source-collaborator-project-mentor ======================================================================== --- title: "Project Manager at Excelerate" company: "Excelerate" period: "April 2023 – Sept 2023" location: "Remote" url: "https://himanshuat.com/experience/excelerate-project-manager" --- # Project Manager — Excelerate *April 2023 – Sept 2023 · Remote* ## About the company Excelerate runs global education and experience programmes for students. I led delivery on one of their events. ## What I did - Led a team of six on a global education event with a $30,000 budget. - Owned the project management docs, including the RACI matrix and risk register. - Delivered on schedule. ## The longer version I ran a team of six on a global education event with a thirty thousand dollar budget, and owned the project management documentation: the RACI matrix, the risk register, the unfashionable artefacts that stop a project quietly falling apart. We delivered on schedule. That is the whole story, and it is the only role here where the thing I learned was administrative rather than technical: how much of delivery is deciding who owns what, in writing, before anybody needs to ask. ## Stack TypeScript, React, Next.js Company site: https://4excelerate.org --- Source: https://himanshuat.com/experience/excelerate-project-manager ======================================================================== --- title: "It's the Fault Policy, Not the Page Size: Why mmap Loses to read() for Out-of-Core Scans on Apple Silicon" category: "Systems" status: "ongoing" url: "https://himanshuat.com/research/mmap-vs-read-apple-silicon" --- # It's the Fault Policy, Not the Page Size: Why mmap Loses to read() for Out-of-Core Scans on Apple Silicon A fault-attributed re-measurement of the CIDR'22 result that mmap underperforms read() for large scans, on 16 KiB-page Apple Silicon, which is absent from that literature. Uses a real 22 GB, 554-million-row financial tick CSV with the parser held constant so the experiment measures the I/O strategy rather than parser cleverness. ## Abstract Memory-mapping a file and letting the kernel handle paging is a common way to scan large datasets, but Crotty et al. (CIDR'22) argued mmap is a poor choice for database scans. That argument was made at the level of mechanism rather than measurement, predates several years of Linux virtual-memory evolution, and was conducted on x86 with 4 KiB pages. Apple Silicon uses 16 KiB base pages and a different memory-management stack, and is essentially unstudied. I re-measure the canonical result on Apple Silicon using a fixed-schema workload: parsing a 22 GB, 554-million-row financial tick CSV. ## Method The parser is a hand-rolled, validating, zero-copy byte scanner held constant across every I/O arm, and sustains around 11 GB/s when isolated from I/O, comfortably faster than the disk. Whatever separates the strategies is therefore I/O rather than compute. Arms cover buffered read(), mmap, mmap with MADV_SEQUENTIAL, mmap with MADV_WILLNEED, and a parallel mmap stream, plus a csv+serde baseline. Each arm runs in a fresh subprocess for clean getrusage peak RSS and minor/major fault counters, with warmup discarded and a working-set sweep via newline-aligned byte limits. ## Results Sweeping the working-set size produces a clean crossover: mmap is 1.2 to 1.3 times faster than buffered read() while the working set fits in RAM, and 0.78 times slower once it exceeds it, with the crossover falling at roughly one-third to two-thirds of physical memory. The slowdown is attributable per fault. Buffered read() takes essentially zero disk faults at any size, whereas mmap's disk faults climb to 1.34 million, one 16 KiB page at a time. The surprise is cross-OS: 4 KiB Linux takes roughly 4 times fewer minor faults than 16 KiB macOS for the same scan, because Linux populates 64 KiB per fault via fault-around and folds anonymous memory into huge pages. x86-64 and AArch64 Linux agree, so the effect is the operating system's fault-population policy, not the page size and not the ISA. madvise cannot recover the gap on macOS: MADV_SEQUENTIAL gives about 15% but leaves the fault count unchanged, and MADV_WILLNEED doubles major faults and is the slowest arm. ## Status In progress. The measurements above are complete on one Apple Silicon machine with one NVMe device and one dataset schema. The cross-OS comparison runs under virtualized and emulated Linux, so I take only fault counts from it, which are architectural, and never timing. Bare-metal Linux confirmation and a second storage device are the outstanding experiments before this is submittable. Targeting arXiv (cs.OS, cs.PF, cs.DB), then a workshop such as DaMoN or HotStorage. Tags: Systems, Operating Systems, Performance, Rust, Databases, Benchmarking --- Source: https://himanshuat.com/research/mmap-vs-read-apple-silicon ======================================================================== --- title: "Conflict-Aware Adapter Composition (C-AAC)" category: "Computer Vision" status: "completed" url: "https://himanshuat.com/research/c-aac" --- # Conflict-Aware Adapter Composition (C-AAC) Authors: Himanshu Published: N/A AuraX's proprietary paradigm shift in Virtual Try-On (VTON) and product imaging. Employs a modular, domain-specific AI architecture using a Conflict-Aware Adapter Composition (C-AAC) algorithm to merge multiple LoRAs on top of the FLUX foundational model without catastrophic forgetting or feature interference. ## Abstract At AuraX, I pioneered a paradigm shift in Virtual Try-On (VTO) by replacing slow, expensive monolithic AI models with a modular, domain-specific AI infrastructure layer tailored for commerce. To generate high-fidelity base images, we utilized the FLUX foundational model, refined via my proprietary Conflict-Aware Adapter Composition (C-AAC) method. C-AAC acts as a three-step conflict resolver that fine-tunes human-set priors, successfully mathematically fusing multiple LoRA adapters (e.g., specific demographics, poses, lighting) into a single optimized runtime model without characteristic feature degradation. ## MultiStageDataPipeline Public datasets were rigorously filtered using my curated multi-stage pipeline focusing on resolution, sharpness, and aspect ratios. To align with commercial e-commerce standards rather than generic 'moody' aesthetics, we trained a proprietary Brand-Centric Aesthetic Model using a human-in-the-loop scoring process. Our initial specialized adapters were fine-tuned on a highly curated corpus of just 5,000 images, forming a hybrid library that achieves production-grade realism across key garment layers like denim, silk, and leather. ## DefensibilityAndResults The C-AAC algorithm's ability to preserve specialization during model merges yields a powerful technical moat, validated quantitatively by high LPIPS scores and a low interference index. Evaluated against base Flux-dev, Google's Imagen 4, OpenAI's ChatGPT (August), and Nano-banana, my AuraX-V1 model demonstrated unequivocal superiority in rendering commercial realism and brand-desirable aesthetics with clean compositions and natural skin textures. Tags: AI, Scene Generation, Diffusion Models, LoRA --- Source: https://himanshuat.com/research/c-aac ======================================================================== --- title: "Flux-VTON+: Hybrid Flux Inpainting and Multi-LoRA Expert Fusion" category: "Generative AI" status: "completed" url: "https://himanshuat.com/research/renu-virtual-try-on" --- # Flux-VTON+: Hybrid Flux Inpainting and Multi-LoRA Expert Fusion Authors: Himanshu Published: N/A A comprehensive framework designed to close the 'Complexity Gap' in Virtual Try-On (VTON) for diverse global garments. Utilizing a dual-stream approach combining Flux Fill for inpainting and Flux Redux for structural style transfer, with a Multi-LoRA Expert Fusion strategy. ## Abstract Virtual Try-On (VTON) has long suffered from a 'Complexity Gap' when applied to non-Western or structurally complex garments like Sarees and Kimonos. I designed the Flux-VTON+ framework to address this gap by synergizing the Flux Fill diffusion architecture with Flux Redux structural guidance. This comprehensive pipeline solves persistent challenges in high-fidelity garment transfer via a Multi-LoRA Expert Fusion strategy that injects specialized knowledge of draping physics and complex occlusion handling directly into the diffusion process. ## ArchitecturalInnovations I architected the pipeline to utilize Segment Anything Model 2 (SAM2) for precision masking and Flux Redux for structural style injection. A core innovation was my Multi-LoRA Expert Fusion technique, which mathematically merges a Draping Physics LoRA and an Occlusion & Depth LoRA into the base UNet. To overcome memory limitations while maintaining fidelity on global garments, I also developed a dynamic Context Window mechanism, ensuring high-resolution processing without exceeding typical VRAM constraints. ## BenchmarkSuperiority In rigorous technical evaluations against baseline SDXL and base Flux models, Flux-VTON+ achieved a state-of-the-art FID of 18.5 and an SSIM of 0.85 on complex Global garments (up from 0.45 for SDXL). The system proved highly successful at handling complex draping physics (e.g., Saree pleats) and occlusions (e.g., hands crossing the garment), dramatically outperforming industry competitors like Alphabake and FashnAI in VTON benchmarks with an 85% overall success rate. Tags: Virtual Try-On, Diffusion Models, Flux, LoRA, Fashion Tech --- Source: https://himanshuat.com/research/renu-virtual-try-on ======================================================================== --- title: "MoE: Brain-Inspired LLM Architecture" category: "Artificial Intelligence" status: "ongoing" url: "https://himanshuat.com/research/moe-brain" --- # MoE: Brain-Inspired LLM Architecture Authors: Himanshu A self-directed research project advancing the thesis that the Mixture of Experts (MoE) architecture is a compelling computational analogue to the human brain's principle of functional specialization. The work deconstructs the neuroscientific foundations of brain organization, provides a technical analysis of MoE models, and synthesizes these domains into a novel, brain-inspired hierarchical MoE architecture. ## Abstract This research advances the thesis that the Mixture of Experts (MoE) architecture, developed for scaling LLMs, represents a compelling computational analogue to the brain's principle of functional specialization[cite: 294]. [cite_start]The core tenets of MoE—modularity, sparse activation, and hierarchical processing—are presented as echoes of a biological blueprint honed by evolution under metabolic constraints[cite: 295, 296, 297]. [cite_start]The work synthesizes neuroscience and AI principles into a novel architecture and critically examines the limitations of the analogy[cite: 299, 300]. ## Hypothesis The Mixture of Experts (MoE) architecture, while developed to solve the problem of scaling LLMs, is one of the most compelling computational analogues to the brain's principle of functional specialization to date[cite: 294]. [cite_start]Its design principles mirror the triad of modularity, hierarchy, and sparsity that underpins the efficiency of biological cognition[cite: 352]. ## ProposedArchitecture A key contribution is the proposal of a Brain-Inspired Hierarchical Mixture of Experts (BI-HME) architecture[cite: 399]. [cite_start]This model features: 1) A multi-level hierarchy analogous to the brain's sensory, association, and prefrontal cortices[cite: 411]. [cite_start]2) Both shared experts for domain-general knowledge and specialized experts for specific tasks[cite: 420]. [cite_start]3) An innovative 'Reliability-Based Gating' mechanism where routing decisions are based on an expert's historical performance, not just input features[cite: 422, 424]. ## Challenges The whitepaper identifies critical gaps between current MoE models and biological reality: 1) **Static Experts vs. Neuroplasticity**: MoE experts are static after training, unlike the brain’s constant, lifelong rewiring[cite: 445, 448]. [cite_start]2) **Oversimplified Gating vs. Cognitive Control**: MoE routers are simple reflexes, unlike the proactive, goal-directed control system of the prefrontal cortex[cite: 453, 456]. [cite_start]3) **Isolated vs. Collaborative Networks**: MoE experts work in parallel isolation, contrasting with the brain's deeply interactive and collaborative network[cite: 463, 468]. Tags: MoE, Cognitive Modeling, LLM, Neuroscience, Self-Research --- Source: https://himanshuat.com/research/moe-brain ======================================================================== --- title: "High-Dimensional Lattice-Based Quantum Encryption (LWE)" category: "Cryptography" status: "stopped" url: "https://himanshuat.com/research/lattice-based-quantum-encryption" --- # High-Dimensional Lattice-Based Quantum Encryption (LWE) Authors: Himanshu This research presents a quantum-resistant encryption algorithm that leverages a high-dimensional lattice framework (n=1000) based on the Learning With Errors (LWE) problem[cite: 290, 302]. [cite_start]The scheme is designed to counter threats from quantum algorithms like Shor's and Grover's, which undermine classical cryptosystems such as RSA and AES[cite: 288, 296]. [cite_start]The parameters are carefully chosen to align with NIST's post-quantum security standards (Category 5, 256-bit quantum security)[cite: 292, 302]. ## Abstract This work details a post-quantum cryptographic (PQC) algorithm founded on the hardness of the Learning With Errors (LWE) problem[cite: 290, 301]. [cite_start]It operates in a 1000-dimensional space with a modulus of q ≈ 2^32 and a discrete Gaussian error distribution, providing robust quantum resistance[cite: 291, 329, 330, 331]. [cite_start]The paper specifies the mathematical operations for key generation, encryption, and decryption, with security guarantees based on worst-case hardness assumptions[cite: 292, 301]. ## MathematicalBackground The scheme's security is based on the LWE problem[cite: 322]. [cite_start]Key generation involves creating a public key (A, b) where b = As + e (mod q) and a secret key s[cite: 333, 337]. [cite_start]For a message m, encryption produces a ciphertext (u, v) by computing u = A^T * r (mod q) and v = b^T * r + floor(q/2)m (mod q), using a random vector r[cite: 341, 343, 344]. [cite_start]Decryption recovers the message m by calculating m' = round(2/q * (v - s^T * u)) (mod 2)[cite: 348, 351]. [cite_start]The paper also discusses optimizations using Ring-LWE and enhanced error reconciliation techniques[cite: 388, 392]. ## ReasonForStopping Shifted focus to more practical AI applications and generative models. ## FutureWork The research identified several avenues for future work, including algorithmic optimizations via parallel processing and hardware acceleration (GPUs/FPGAs)[cite: 448]. [cite_start]It also proposed investigating advanced error reconciliation techniques and conducting field testing in real-world environments like secure cloud storage and IoT systems to evaluate practical performance[cite: 450, 452]. Tags: Quantum Cryptography, Lattice Theory, Post-Quantum, Security, LWE --- Source: https://himanshuat.com/research/lattice-based-quantum-encryption ======================================================================== --- title: "GPT-Neo: Transformer Implementation" category: "Natural Language Processing" status: "completed" url: "https://himanshuat.com/research/gpt-neo" --- # GPT-Neo: Transformer Implementation Authors: Himanshu Published: N/A Complete implementation of the 'Attention is All You Need' paper, building a GPT-style language model from scratch with detailed documentation and a training pipeline. ## Abstract A comprehensive, from-scratch implementation of the transformer architecture as described in the seminal paper 'Attention is All You Need'. This project focuses on building a decoder-only, GPT-style model. ## Implementation Built using PyTorch, the model includes multi-head self-attention, positional encoding, feed-forward networks, and layer normalization. The repository also contains a complete training and inference pipeline. ## Results The model was successfully trained on various text corpora, demonstrating its ability to generate coherent text and understand context. It serves as a strong educational baseline for transformer architectures. ## Learnings Gained a deep, practical understanding of attention mechanisms, model architecture, and the challenges of training large language models, including managing computational resources and preventing overfitting. Tags: Transformers, NLP, Deep Learning, PyTorch, Implementation Paper: https://github.com/Himasnhu-AT/gpt-neo --- Source: https://himanshuat.com/research/gpt-neo ======================================================================== --- title: "Custom Neural Network for MNIST" category: "Machine Learning" status: "completed" url: "https://himanshuat.com/research/custom-n-n-mnist" --- # Custom Neural Network for MNIST Authors: Himanshu Published: N/A A hand-crafted neural network built from the ground up in Python with NumPy for MNIST digit recognition. This project was undertaken to demonstrate and solidify an understanding of the fundamental concepts of machine learning without relying on high-level frameworks like TensorFlow or PyTorch. ## Abstract This project details the implementation of a custom neural network for digit recognition on the classic MNIST dataset. The entire network, including layers, activation functions, and the backpropagation algorithm, was built using only the NumPy library. ## Approach The core objective was to build a functional neural network without high-level ML libraries. This required a deep dive into the mathematics of forward and backward passes, implementing the backpropagation algorithm by hand, and managing weights and biases manually. ## Results After tuning hyperparameters such as learning rate and the number of hidden neurons, the custom-built network achieved a respectable 85% accuracy on the MNIST test set, demonstrating the viability of the from-scratch approach. ## Educational This serves as an excellent learning project for anyone seeking to understand the fundamental mechanics of neural networks. It provides a clear, practical insight into how data flows through a network and how learning occurs via gradient descent and backpropagation. Tags: Neural Networks, MNIST, From Scratch, NumPy, Educational Paper: https://github.com/Himasnhu-AT/Custom-N_N-MNIST --- Source: https://himanshuat.com/research/custom-n-n-mnist