# Build on Anthropic: Full Content > This file contains the full text of every page in the wiki. > Generated automatically at build time. --- # Case Studies URL: https://buildonanthropic.com/case-studies # Case Studies *Anthropic publishes 256 customer stories and roughly 95 more that never appear in its own index. This is the organized view, cut so you can find the company that looks like yours.* ![Hand-drawn line illustration in warm terracotta and cream: two hands fan out five blank story cards like a card hand, studied by an abstract face drawn in a single line.](/img/pages/case-studies.jpg) --- ## Why this section exists `claude.com/customers` is a good corpus behind a bad index. The page lazy-loads a small batch of stories at a time, so the visible count understates the real one by roughly 90 percent. The sitemap is authoritative: there are **256 stories**. There is also a second corpus almost nobody sees. Around **95 additional customer stories** live outside the index entirely, spread across news posts, blog posts, and conference sessions. Several major named customers have no `/customers` page at all, and a few companies publish stronger numbers off-index than on their own case-study page. Everything below is the company's own published figure, quoted briefly and linked to the source. ## Start with the cut that matches your question **The master list is [The Index](/case-studies/the-index): all 256 published stories, one row each, alphabetical and searchable, every row linked to the official page.** | If you are asking | Go here | |---|---| | "Is there a story about company X?" | [The Index](/case-studies/the-index) | | "Has a company at my stage done this?" | [The startup cut](/case-studies/startups) | | "What does the strongest evidence actually look like?" | [Strength of evidence](/case-studies/strength-of-evidence) | ## What the corpus looks like at a glance **By industry.** Software and SaaS dominates at 125 stories, roughly half the corpus. Then financial services (20), professional services (17), beneficial deployments and nonprofits (15), education (14), cybersecurity (10), healthcare (8, plus 5 in life sciences), and legal (8). Cybersecurity is worth noting out of proportion to its size: it carries the highest average metric quality in the corpus. Security teams publish falsifiable numbers because their buyers demand them. **By use-case pattern.** Coding and developer productivity is the largest at roughly 45 stories, followed by agents that do a whole job end-to-end (~40), document intelligence and extraction (~35), search and retrieval (~25), content generation (~25), classification and triage (~20), customer support automation (~20), voice (~10), and computer use or browser automation (~8). The distribution is itself a finding. Coding is the wedge most companies enter through, but the end-to-end agent category is close behind and is where the more interesting business outcomes cluster. ## How to read the numbers responsibly Not all published metrics are equally strong, and treating them as interchangeable is how you lose credibility with a technical audience. - **Controlled or benchmarked** results, where a method is disclosed, are the ones to lead with. - **Hard operational numbers** are self-reported but specific and falsifiable. - **Directional numbers** are credible but have no disclosed method. Use them for color, not for an argument. A handful of published pages also contradict themselves, print another company's statistics, or explicitly label their own figures illustrative. Those are flagged individually in [Strength of evidence](/case-studies/strength-of-evidence). Check a number there before you put it in a deck. ## Further Reading - [The Index](/case-studies/the-index) all 256 published stories in one searchable table - [The startup cut](/case-studies/startups) the companies that were small when they started - [Strength of evidence](/case-studies/strength-of-evidence) which numbers hold up, and which to avoid - [Reference Index](/reference) every published source, dated --- # The Startup Cut URL: https://buildonanthropic.com/case-studies/startups # The Startup Cut *Twenty-one of Anthropic's published customer stories are Y Combinator companies, and they cover every rung of the ladder from current batch to acquisition.* ![Hand-drawn line illustration in warm terracotta and cream: a single hand places the first small cream block at the edge of a wide, nearly empty foundation slab, watched by an abstract one-line face.](/img/pages/startups.jpg) --- ## Why this cut matters most Enterprise case studies are easy to dismiss if you are three people with no revenue. The useful question is not "does Claude work at Rakuten," it is "has a company at my size and stage shipped something real on it." Anthropic explicitly tags **29 stories** as startups. Roughly 25 more are tagged small and describe themselves as startups in the body. Separately, **21 of the 256 stories are Y Combinator companies**. ## The YC ladder, told entirely with published customers This is the single most useful narrative in the corpus for an early-stage audience, because every rung has a real story attached. | Stage | Companies | |---|---| | **Current or recent batch** | Juno, Tasklet, Crunched, cubic | | **Series A** | OffDeal, Greptile | | **Series B** | Emergent, Campfire, Mintlify, Vapi, Mutiny | | **Series D and beyond** | Legora, Vanta, Benchling, Replit, Apollo | | **Exit** | Newfront, acquired by WTW | A few worth knowing by name: - **Juno**, a current-batch company, reports reaching 100,000 patients since an October launch, with the app built and maintained by a two-person team using Claude Code. ([source](https://claude.com/customers/juno)) - **Tasklet**, also current-batch, reports 160 percent month-over-month revenue growth to over $2.5M ARR within five months of launch, and 450,000 agent actions executed per day, from a five-engineer team. ([source](https://claude.com/customers/tasklet)) - **cubic** reports a first code review delivered in 2 minutes instead of 2 hours, and that 98 percent of AI-generated comments receive positive developer feedback. ([source](https://claude.com/customers/cubic)) - **Emergent** reports $25M ARR in 4.5 months of commercial launch with more than 2 million users, after three pivots. ([source](https://claude.com/customers/emergent)) - **OffDeal** reports raising eval accuracy from 25 percent to 85 percent after moving to the Claude Agent SDK and iterating on tool design and prompting, noting that the switch alone moved the score from 25 to 60. ([source](https://claude.com/customers/offdeal)) - **Greptile** reports roughly 90 percent cache hit rates and a million issues caught per month with a team of ten engineers. ([source](https://claude.com/customers/greptile)) ## Named YC batch membership is rare in the source material Worth stating plainly for anyone rebuilding this list: the case studies themselves name Y Combinator only three times. The mapping above was built by checking the live YC directory against each customer's actual outbound domain. That verification is not optional, because the directory is full of name collisions. There is a YC company called Warp, one called Clay, one called Assembled, one called Athena, and one called Artemis, and **in each case it is a different company from the Anthropic customer with the same name.** Matching on name alone produces a list that is roughly a third wrong. ## What the startup stories have in common Reading the startup subset as a group, a few patterns repeat often enough to be worth copying. **Small teams shipping disproportionate scope.** Two people and 100,000 patients. Five engineers and $2.5M ARR. Ten engineers catching a million issues a month. The recurring shape is a team that would have needed to be four times larger a few years ago. **Speed from adoption to production measured in days.** Several report going from decision to deployed in under a week. One reports a production deployment within days of adopting the Agent SDK. **Evals as the thing that actually moved the number.** The most credible startup stories do not credit the model alone. They credit iterating against an eval set, and they publish the before and after. **Unit economics treated as an engineering problem.** The startups with the best cost stories talk about cache hit rates and prompt structure, not about picking a cheaper model. One reports cutting daily model spend meaningfully by raising the share of input tokens served from cache. ## A note on what this list is not Funding stages and batch labels here come from secondary sources and are best-effort. The Claude-related claims come from the companies' own published pages and are quoted briefly with links. Where a company published a figure without disclosing a method, treat it as directional. See [Strength of evidence](/case-studies/strength-of-evidence). ## Further Reading - [Case Studies overview](/case-studies) the full corpus and how it is cut - [Strength of evidence](/case-studies/strength-of-evidence) which numbers hold up - [Reference Index](/reference) every published source, dated --- # Strength of Evidence URL: https://buildonanthropic.com/case-studies/strength-of-evidence # Strength of Evidence *Every number in the corpus is first-party. That does not make them equally strong. Lead with the ones that disclose a method.* ![Hand-drawn line illustration in warm terracotta and cream: a balance scale where one solid cream block outweighs a pan of loose paper scraps, steadied by a hand and examined by an abstract one-line face.](/img/pages/evidence.jpg) --- ## Why tier at all Quoting a directional productivity percentage to a skeptical engineer costs you the room. Quoting a benchmark with a disclosed method and a control group wins it. The difference is not the size of the number, it is whether the reader can tell how it was produced. One useful fact about this corpus first: **every figure comes from Anthropic's own customer-story pages**, which are first-party company statements published with named attribution. No press coverage feeds it, so there is no journalist-versus-company reliability question to adjudicate. The distinctions below are entirely about method. ## Tier A: controlled, benchmarked, or audited Quote these first. Each discloses how the number was produced. - **Smartsheet** is the strongest structurally, because it is the only story in the corpus with a same-team control group: engineers using Claude Code compared against peers on the same teams. ([source](https://claude.com/customers/smartsheet)) - **Freedom Forever** published a real head-to-head with a stated methodology, building eight mock sites at five iterations each across 40 runs. ([source](https://claude.com/customers/freedom-forever)) - **eSentire** evaluated against senior human analysts in production across more than 500 adjudicated outcomes. ([source](https://claude.com/customers/esentire)) - **Rising Academies** is the only externally validated result in the corpus, reporting an effect size from school-based studies run by researchers at Oxford and JPAL. ([source](https://claude.com/customers/rising-academies)) - **Apollo** ran blind testing and then a production A/B, attributing a retention change to the model swap alone. ([source](https://claude.com/customers/apollo)) - **Semgrep**, **Graphite**, **Descript**, **Shortcut**, **Wordsmith**, **GC AI**, and **Satispay** each disclose an evaluation set, a test-case count, or a structured comparison window. ## Tier B: hard operational numbers Self-reported, no disclosed method, but specific and falsifiable enough to be useful. These are counts and durations rather than percentages, which is what makes them checkable in principle. Examples include Wiz on a line-count migration and the hours it took, Stripe on a language migration across a stated engineer count, Rakuten on time-to-market and an autonomous run length, LG CNS on APIs converted out of a total, Spotify and Delivery Hero on merged pull requests, Novo Nordisk on documentation turnaround, Dust on daily model spend, and Notion on cost and latency reduction from caching. ## Tier C: credible directional numbers The bulk of the corpus. Percentage productivity gains, time savings, adoption rates, none with a disclosed method. Useful for color and for showing a direction of travel. Always attribute them as what the company reports rather than as measured fact. ## Tier D: use with care **Metrics the company itself flags as illustrative.** At least one page prints a disclaimer stating its own quoted statistics are illustrative only and that results vary by configuration and context. Take the company at its word and do not quote those numbers as outcomes. **Projections rather than results.** Several figures are forward-looking: revenue projected under an assumed monetization model, time savings a company expects rather than has recorded, hours it is on track to redirect. These are plans. Label them as plans. **Stories with no quantitative metrics at all.** Roughly fourteen stories are qualitative only. They are still useful as existence proofs that a category of company shipped something, and useless as evidence of magnitude. ## Pages that contradict themselves Seven pages carry internal inconsistencies or cross-page contamination. Do not smooth these over. Pick one figure and cite the exact sentence it came from. | Page | The problem | |---|---| | **Brex** | Prints two different expense-automation rates, and two different monthly hours-saved figures, on the same page. | | **Replit** | Headline and body disagree on user count, and the ARR sentence appears in two incompatible forms. | | **L'Oréal** | Accuracy stated one way in the bullet list and another in the outcome section. | | **Qualified Health** | Headline callout and body bullet disagree on patient population size. | | **Graphite** | The top-of-page stat callouts belong to a different company's story entirely and match nothing in the body. | | **Emergent** | Its callout block carries an injected card containing another company's compliance metric. | | **Zoom** | A satisfaction figure appears on the page that belongs to a different customer's story. | The Graphite, Emergent, and Zoom cases are the same underlying defect: callout contamination across pages. If a headline statistic does not appear anywhere in the body prose, do not trust it. ## One more attribution boundary Several customer pages carry Anthropic marketing cards inside the content flow, making claims about enterprise AI adoption or time-to-production. Those are Anthropic's claims, not the customer's. When you quote a customer, quote the customer. ## Further Reading - [Case Studies overview](/case-studies) the full corpus and how it is cut - [The startup cut](/case-studies/startups) companies that were small when they started - [Reference Index](/reference) every published source, dated --- # The Index URL: https://buildonanthropic.com/case-studies/the-index # The Index *Every published Anthropic customer story. One row per story, 256 rows, alphabetical. Each row links to the official page, which is always the authority.* ![Hand-drawn line illustration in warm terracotta and cream: a hand's finger touches one tile in a large orderly grid of small cream tiles, scanned by an abstract one-line face.](/img/pages/the-index.jpg) --- ## How to search this page - **Wiki search.** The search box at the top of this site indexes every row. - **Browser find.** The whole index is one page, so Ctrl+F / Cmd+F reaches everything: a company name, a vertical like "legal", a surface like "Claude Code", or a number. - **For agents.** This wiki publishes [llms-full.txt](/llms-full.txt), which carries this entire table in plain text. - **The source wins.** Every company name links to its official claude.com/customers page. If a number matters to a decision, click through and read the original. Column notes: **Vertical** is normalized to a short controlled vocabulary. **Surface** compresses the Anthropic products the story names. **Standout number** is one metric quoted from the story, shortened but never reworded; "qualitative" means the story publishes no usable number. Seven stories carry self-contradicting figures; their cells avoid the contested numbers or point to [Strength of evidence](/case-studies/strength-of-evidence). **YC** marks companies confirmed in the Y Combinator directory (21 of 256 rows). A few companies appear twice because Anthropic published two distinct stories about them; each row is one story. ## A to M | Company | What they built | Vertical | Surface | Standout number | YC | |---|---|---|---|---|---| | [Adalat AI](https://claude.com/customers/adalat-ai) | WhatsApp helpline translating court orders for litigants | legal | API, Claude Code | 50–200% productivity gains across engineering teams with Claude Code | | | [Advantage Solutions](https://claude.com/customers/advantage-solutions) | Compliance, finance, and comms automation for retail ops | retail | Enterprise, Claude Code, Cowork | Compressed a recurring weekly 10-hour finance forecast validation workflow to under 30 minutes | | | [Advolve](https://claude.com/customers/advolve) | Agent that runs digital ad campaigns end to end | sales & marketing | API | 90% reduction in operational work time | | | [AES](https://claude.com/customers/aes) | Multi-agent safety audit report generation | other | API, Vertex | 96% faster audit results generation | | | [AirOps](https://claude.com/customers/airops) | Marketing content agents that draft and quality-check articles | sales & marketing | API, Claude Code, Enterprise | Time to quality went from 60 hours to 5 hours | | | [Airtree](https://claude.com/customers/airtree) | Firmwide Cowork skills for board prep and research | fintech | Cowork, Claude Code, MCP | market and competitor research that would have taken two days can now be done in minutes | | | [Alexa+](https://claude.com/customers/alexa-plus) | Claude models power the Alexa+ assistant | other | API, Bedrock | qualitative | | | [Amazon Connect / Amazon Q in Connect](https://claude.com/customers/amazon-q-in-connect) | Real-time agent assist and self-service in contact centers | customer support | API | 61 languages supported without additional training | | | [Amira Learning](https://claude.com/customers/amira) | Voice reading tutor for elementary students | education | API | 10 billion words read aloud and analyzed worldwide | | | [Anything](https://claude.com/customers/anything) | Agent that builds and ships full-stack apps from prompts | dev tools | API, Agent SDK, Claude Code | 800,000+ apps created in five months | | | [Apollo](https://claude.com/customers/apollo) | Personalized outbound sales messaging at scale | sales & marketing | API | 35% boost in meeting bookings with Claude messaging | ✓ | | [AppFolio](https://claude.com/customers/appfolio) | AI responses for property manager communications | other | API, Bedrock | 3x adoption rates due to Claude 3 Haiku's faster response times | | | [Armanino](https://claude.com/customers/armanino) | Auto-drafted audit document follow-up notes | consulting | API, Bedrock, Vertex | 65% reduction in time spent on manual writing tasks | | | [Artemis](https://claude.com/customers/artemis) | Agents that build detections and run investigations | security | API, Bedrock, Claude Code | 90% increase in detection coverage with environment-specific detections built and tuned from day one | | | [Asana](https://claude.com/customers/asana) | Asana AI workflows, chat, and summaries | productivity | API | 10x faster insights data-driven decision making with Asana AI | | | [Asana (AI Teammates Q&A)](https://claude.com/customers/asana-qa) | AI Teammates that pick up tasks inside projects | productivity | Managed Agents | What used to take days of back-and-forth becomes a fifteen-minute review cycle | | | [ASAPP](https://claude.com/customers/asapp) | GenerativeAgent automated customer service | customer support | API, Bedrock | 25-40% improvement in core business metrics vs other AI models | | | [Assembled](https://claude.com/customers/assembled) | Ticket triage, routing, and response drafting | customer support | API | 20% increase in customer satisfaction while decreasing support spend | | | [Athena Intelligence](https://claude.com/customers/athena) | Autonomous analyst workflows for regulated enterprises | productivity | API | 4–8 weeks to production vs. 6–12 month industry standard | | | [Attention](https://claude.com/customers/attention) | Call scoring, CRM autofill, and follow-up drafting | sales & marketing | API | 1.6 million hours of admin work automated freeing sales reps to focus on closing deals | | | [Audience Strategies](https://claude.com/customers/audience-strategies) | Interview synthesis for industry research reports | consulting | Enterprise | 3 months to 4 weeks report creation timeline compression | | | [Augment Code](https://claude.com/customers/augment-code) | Codebase intelligence for enterprise developers | dev tools | API, Vertex | 4-8 months to 2 weeks project timeline acceleration | | | [Aura Intelligence](https://claude.com/customers/aura) | Job title classification across 200 million records | HR & recruiting | API, Bedrock | 94% accuracy in overall classification tasks | | | [Banner Health](https://claude.com/customers/banner-health) | Oncology chart prep from patient records | healthcare | API, Bedrock, Claude Code | 85% of users report time savings with improved work accuracy | | | [Base44](https://claude.com/customers/base44) | No-code AI app builder | dev tools | API | Base44 has grown to more than 10% of the AI-powered app builder market within months. | | | [Benchling](https://claude.com/customers/benchling) | Lab data assistants for structuring and SQL queries | healthcare | API, Bedrock | 2 weeks saved transforming complex data with Data Entry Assistant | ✓ | | [Binti](https://claude.com/customers/binti) | Drafts home-visit reports and answers over case files | nonprofit | API | 50% faster home visit report writing time, from 3-4 hours to under 2 hours | | | [Biomni](https://claude.com/customers/biomni) | Autonomous biomedical research agent | healthcare | API | 800x faster 35 minutes vs 3 weeks for analysis | | | [Bito](https://claude.com/customers/bito) | AI code review and coding agents | dev tools | API | 89% reduction in pull request cycles | | | [Blank Metal](https://claude.com/customers/blank-metal-qa) | Runs its consulting ops on Cowork and Claude Code | consulting | Cowork, Claude Code, Enterprise | +700 people trained in Claude Code and Cowork | | | [Block](https://claude.com/customers/block) | Internal agent goose for SQL, code, and automation | fintech | API, MCP | 75% of engineers saving 8 to 10+ hours every week using codename goose | | | [BlueFlame](https://claude.com/customers/blueflame) | Private equity deal-room document analysis | fintech | API | Reduced document analysis time from 4+ hours to minutes | | | [Bluenote](https://claude.com/customers/bluenote) | Regulatory document drafting for life sciences | healthcare | API | 10x faster analysis for scientists parsing complex protocols | | | [Bolt / StackBlitz](https://claude.com/customers/bolt) | Design-system agent for on-brand app prototypes | dev tools | Agent SDK, Claude Code | 10,000+ users uploading their own design systems to Bolt | | | [Box](https://claude.com/customers/box) | Box Agent generates documents from stored content | productivity | API | Redlined a contract in 2 minutes instead of an afternoon of manual review | | | [Brainlabs](https://claude.com/customers/brainlabs) | Employee-authored Skills automate agency workflows | sales & marketing | Cowork, Enterprise, Managed Agents | ~400 Skills authored by employees in four weeks | | | [Braintrust](https://claude.com/customers/braintrust) | AI screening interviews and job description generation | HR & recruiting | API, Bedrock | 25% more applicants for Claude-powered job descriptions | | | [Brand.ai](https://claude.com/customers/brand-ai) | Keeps global content on-brand automatically | sales & marketing | API | Enables one copywriter to manage 600 pieces of content | | | [Brex](https://claude.com/customers/brex) | Expense categorization and 100% transaction auditing | fintech | API, Bedrock | 94% compliance rate vs 70% industry standard | | | [Brian Impact Foundation](https://claude.com/customers/brian-impact-foundation) | Candidate research system for fellowship selection | nonprofit | API | 100x more candidates reviewed with Claude - 20,000 vs 200 | | | [Bubble](https://claude.com/customers/bubble) | AI app generator and in-editor editing agent | dev tools | API, Claude Code | 2x first-week user activation rate after launching the Claude-powered AI app generator | | | [bunq](https://claude.com/customers/bunq) | Finn banking assistant and 5-minute onboarding | fintech | Claude Code | Account opening that once took days with paperwork and branch visits now takes 5 minutes. | | | [Campfire](https://claude.com/customers/campfire) | Conversational financial reporting and reconciliation | fintech | API | 3-day reduction in monthly close time | ✓ | | [Canva](https://claude.com/customers/canva) | Company-wide Claude for design and code work | media | Enterprise | 5,000 employees empowered with Claude for Work across all teams | | | [Carta Healthcare](https://claude.com/customers/carta-healthcare) | Clinical registry data abstraction with citations | healthcare | API, Bedrock | Up to 66% reduction in time required for clinical data abstraction | | | [Charm Industrial](https://claude.com/customers/charm-industrial) | Carbon credit verification automation | other | API | 3-6 months to days carbon credit verification time | | | [ChatAndBuild](https://claude.com/customers/chatandbuild) | Natural-language AI app building in 40+ languages | dev tools | API | 140,000 users across 111 countries | | | [Chatbase](https://claude.com/customers/chatbase) | Support agents that answer and act in brand voice | customer support | API | Tripled user adoption | | | [ChatPlace](https://claude.com/customers/chatplace) | Instagram content and DM agents for creators | media | API, Claude Code, MCP | 15-20 hours saved per creator per week across content and DMs | | | [Chronograph](https://claude.com/customers/chronograph) | Company-wide Claude Enterprise wired to internal tools | fintech | Enterprise, MCP | Achieves 100% employee adoption with 80% daily active usage across all 150+ team members | | | [Circleback](https://claude.com/customers/circleback) | Meeting notes, action items, and automations | productivity | API | qualitative | ✓ | | [CircleCI](https://claude.com/customers/circleci) | Chunk agent ships validated green pull requests | dev tools | API, Agent SDK, Claude Code | 90% of the engineering team on Claude Code, with daily usage jumping 9x since structured adoption began | | | [ClassDojo](https://claude.com/customers/classdojo) | Teaching assistants for lessons, behavior, and parent comms | education | API | 45 million users kids, families, and teachers in 180 countries | | | [Classmethod](https://claude.com/customers/classmethod) | AI self-reviewing code generation across delivery | consulting | Claude Code, MCP | 80% decrease in time spent reviewing code | | | [Clay](https://claude.com/customers/clay) | Claygent researches leads and personalizes outreach | sales & marketing | API | Hundreds of hours saved with automated data collection | | | [CodeRabbit](https://claude.com/customers/coderabbit) | AI code review on pull requests | dev tools | API | 86% faster code delivery from days to hours | | | [CodeWords / Agemo](https://claude.com/customers/codewords) | Conversational workflow builder that deploys automations | dev tools | API, Agent SDK | 6 minutes to build complex multi-source automations, down from 25 minutes | | | [Cogent](https://claude.com/customers/cogent) | Agents that investigate and remediate vulnerabilities | security | API | 97% reduction in time that critical vulnerabilities remain open | | | [Cognition](https://claude.com/customers/cognition) | Devin, an autonomous ticket-to-PR software engineer | dev tools | API | 3.5x increase in merged PRs per week after adopting Claude Sonnet 3.6 | | | [Coinbase](https://claude.com/customers/coinbase) | Agentic support chatbot and rep assist | fintech | API, Bedrock, Vertex | 35-50 apps built with internal CB-GPT platform | | | [Copy.ai](https://claude.com/customers/copy-ai) | Long-form GTM content generation | sales & marketing | API | 4x increase in content output | | | [Cove](https://claude.com/customers/cove) | Visual canvas workspace driven by two Claude models | productivity | API | 30% faster response times compared to previous solutions | | | [Cox Automotive](https://claude.com/customers/cox-automotive) | Personalized dealer emails and vehicle listings | retail | API, Bedrock | 2x increase in consumer lead responses and test drives | | | [Cox Communications & Accenture](https://claude.com/customers/cox-and-accenture) | Marketing and sales agents for B2B broadband | sales & marketing | API, Claude Code, Cowork | 7x ROI on Cox Communications' first year of AI investment | | | [Cox Communications (Q&A)](https://claude.com/customers/cox-communications-qa) | Multi-agent B2B funnel plus 2,500 employee builders | sales & marketing | Claude Code, Cowork, Enterprise | We're on our way to 2,500 Claude Code users and climbing, more than half of them non-development engineers | | | [CRED](https://claude.com/customers/cred) | Claude Code across the development lifecycle | fintech | API, Claude Code | 2x faster execution speed for feature delivery and fixes | | | [Crunched](https://claude.com/customers/crunched) | Autonomous Excel model building for finance | fintech | API | >50% time savings on Excel modeling for finance and consulting workflows | ✓ | | [cubic](https://claude.com/customers/cubic) | Dependency-aware AI pull request review | dev tools | API | 60x faster feedback cycle - 2 min vs 2 hours | ✓ | | [Cursor (Q&A)](https://claude.com/customers/cursor-qa) | Claude models power Cursor's coding agents | dev tools | Claude Code | 60% of the Fortune 500 builds software with Cursor | | | [Decagon](https://claude.com/customers/decagon) | Support agents that complete tasks end to end | customer support | API | 70% reduction in over-inferencing rates | | | [Delight.ai / Sendbird](https://claude.com/customers/delightai-qa) | Claude Code builds agent debugging and regression tooling | customer support | Claude Code, Bedrock | 1 week → 1–2 days Time to fix and redeploy a production AI agent issue | ✓ | | [Delivery Hero](https://claude.com/customers/delivery-hero) | Herogen agent turns Jira tickets into pull requests | dev tools | API, Claude Code, Vertex | 85% success rate Most tasks completed with zero to one developer interactions | | | [Descript](https://claude.com/customers/descript) | Underlord, an agentic video co-editor | media | API | 13-23% higher intent adherence than leading competitors | | | [Doctolib](https://claude.com/customers/doctolib) | Claude Code automation across the engineering org | healthcare | Claude Code | Migrated their entire visual regression testing tool in hours rather than weeks | | | [Dolly](https://claude.com/customers/dolly) | Grades assessments and writes progress reports | education | API, Bedrock | 60-70% reduction in marking time for digital assessments | | | [Dust](https://claude.com/customers/dust) | No-code enterprise agents wired into company stacks | productivity | API, Claude Code, MCP | $10k/day saved on model spend after optimizing prompt caching | | | [Duvo](https://claude.com/customers/duvo) | Agents that operate SAP and supplier portals | retail | API, Agent SDK, MCP | €2.8M+ in annualized savings captured in three months for one multi-billion-euro retailer | | | [Elation Health](https://claude.com/customers/elation-health) | Cited chart summaries inside the EHR | healthcare | API, Claude Code | 61% reduction in median time to first insight for chart review | | | [Emergent](https://claude.com/customers/emergent) | Autonomous agents that generate full-stack apps | dev tools | API | Achieved $25M ARR in 4.5 months of commercial launch | ✓ | | [Epic Systems](https://claude.com/customers/epic-systems) | Company-wide Claude Code, including non-developers | healthcare | Claude Code, MCP | Over half of Claude Code usage is from non-developers | | | [eSentire](https://claude.com/customers/esentire) | Autonomous multi-agent threat investigations | security | API, Bedrock, Agent SDK | 99.96% ransomware containment before any file encryption, more than double the industry average of 44% | | | [Eve Legal](https://claude.com/customers/eve-legal) | Case file analysis, chronologies, and demand letters | legal | API, Claude Code, Bedrock | 60 days faster settlement times for firms on Eve | | | [Factory](https://claude.com/customers/factory) | Droids for review, refactors, and ticket-to-PR work | dev tools | API | 550,000 hours of development time saved | | | [Figma](https://claude.com/customers/figma) | Prompts and mockups become working prototypes | media | API, Claude Code | qualitative | | | [Fountain](https://claude.com/customers/fountain) | Automated candidate screening and hiring agents | HR & recruiting | API, Claude Code | 50% reduction in manual screening effort | ✓ | | [Freedom Forever](https://claude.com/customers/freedom-forever) | Agents that file solar permits on utility websites | other | API, Agent SDK, Claude Code | Achieved 97.5% success rate on complex web automation benchmarks where alternative frameworks hit 20-60% | | | [FutureHouse](https://claude.com/customers/futurehouse) | Research agents that search, cite, and assess novelty | healthcare | API | literature reviews completed in days instead of months | | | [Gambit Robotics](https://claude.com/customers/gambit) | Recipe parsing and real-time cooking guidance | dev tools | API, Claude Code | Achieves ~97%+ successful recipe parses across their pipeline | | | [Gamma](https://claude.com/customers/gamma) | Generates presentations and websites from prompts | sales & marketing | API | 30% improvement in user satisfaction | | | [Garvan Institute of Medical Research (Q&A)](https://claude.com/customers/garvan-institute-qa) | Claude Code for genomic analysis and research ops | healthcare | Claude Code | 20+ software engineers and data scientists actively using agentic development tools | | | [GC AI](https://claude.com/customers/gc-ai) | Governed legal workflows with cited research | legal | API | 14 hours saved per week on average for in-house legal teams using the platform | | | [Genspark](https://claude.com/customers/genspark) | Super Agent orchestrating 150+ tools | productivity | API, Claude Code | $250M ARR since pivoting to the Super Agent in early 2025 | | | [GitLab](https://claude.com/customers/gitlab) | GitLab Duo AI features | dev tools | API | 25-50% productivity gains across internal workflows with Claude | | | [GitLab (Claude Enterprise)](https://claude.com/customers/gitlab-enterprise) | Claude for Work across internal departments | dev tools | Enterprise | 98% satisfaction reported by GitLab team members | | | [Grab](https://claude.com/customers/grab) | Multilingual merchant assistant and consulting | retail | API | 5.7pp increase in merchant issue resolution rate | | | [Gradial](https://claude.com/customers/gradial) | Content authoring into CMS under brand governance | sales & marketing | API, Cowork, Claude Code | 300+ hours of bulk content operations now take less than 10 hours | | | [Gradient Labs](https://claude.com/customers/gradient-labs) | End-to-end resolution of regulated support queries | customer support | API | 80-90% resolution rates for automated support | | | [Grafana Labs](https://claude.com/customers/grafana) | Assistant that writes queries and builds dashboards | dev tools | API, MCP | Under 8 weeks from prototype to full private preview | | | [Graphite](https://claude.com/customers/graphite) | AI pull request reviewer with one-click fixes | dev tools | API | 40x faster pull request feedback loop, from 1 hour to 90 seconds | | | [Greptile](https://claude.com/customers/greptile) | Multi-hop autonomous code review agent | dev tools | Agent SDK, Claude Code, MCP | 1 million issues caught per month | ✓ | | [Gumroad](https://claude.com/customers/gumroad) | Support staff ship code and features directly | retail | API | 300% increase in new features shipped to production | | | [Harvey](https://claude.com/customers/harvey) | Model choice for legal drafting and review | legal | API | Under 1 month deployment across enterprise platform | | | [Headstart](https://claude.com/customers/headstart) | Claude writes nearly all client code | consulting | API | 90-97% of client code written by Claude | | | [Hebbia](https://claude.com/customers/hebbia) | Matrix analyzes hundreds of documents in parallel | fintech | API | 1/3 of top 50 asset managers served | | | [Hostinger](https://claude.com/customers/hostinger) | Full website generated from one API call | dev tools | API | 3.5 million users served across 150+ countries | | | [HubSpot](https://claude.com/customers/hubspot) | Claude across engineering, marketing, and support | sales & marketing | API, Claude Code, MCP | Up to 40% productivity increase across web development and content creation workflows | | | [HubSpot](https://claude.com/customers/hubspot-qa) | Cowork skills for marketing memos and internal tools | sales & marketing | Cowork, MCP, Claude Code | Multi-week internal builds compressed to three hours | | | [Humach](https://claude.com/customers/humach) | Voice agents and live agent assist | customer support | API, Bedrock | 15-20% increase in operational efficiency | | | [Hume AI](https://claude.com/customers/hume) | Claude speaks through EVI's empathic voice interface | dev tools | API | Over 2 million minutes of AI voice conversations completed | | | [IFS Nexus Black](https://claude.com/customers/ifs) | Resolve diagnoses equipment faults from images and sensors | other | API, MCP, Agent SDK | Resolves faults in assets at a plant, site or in the field 25% faster | | | [IG Group](https://claude.com/customers/ig-group) | Company-wide Claude for HR, marketing, and SQL | fintech | Enterprise | 70 hours saved weekly for analysts with AI-assisted analytics and query handling | | | [Inscribe](https://claude.com/customers/inscribe) | AI risk agents for fraud detection and KYC | fintech | API, Bedrock | Reduce fraud review time by 20x, from 30 minutes to 90 seconds | ✓ | | [Intercom](https://claude.com/customers/intercom) | Fin resolves support queries end to end | customer support | API | 86% resolution rate with human-quality support responses | | | [Intuit](https://claude.com/customers/intuit) | Plain-language tax explanations in TurboTax | fintech | API, Bedrock | Served millions of TurboTax customers during peak season | | | [JAKALA](https://claude.com/customers/jakala) | AI Factory with production agents for client accounts | consulting | Agent SDK, Enterprise, MCP | ~70% reduction in time spent on systematic and repetitive tasks across senior delivery teams | | | [Jamf](https://claude.com/customers/jamf) | Governed company-wide Claude Enterprise rollout | security | Enterprise, Cowork, Bedrock | 89% active usage among licensed employees within eight weeks of company-wide launch | | | [Jamf](https://claude.com/customers/jamf-qa) | Cowork skills turn spreadsheets into guided workflows | security | Cowork, Enterprise, MCP | 45 minutes to build a conversational UI in Cowork | | | [JetBrains](https://claude.com/customers/jetbrains) | Code generation and the Junie coding agent in IDEs | dev tools | API, Bedrock | 59% increase in refactoring success rate | | | [JetBrains](https://claude.com/customers/jetbrains-2) | Native Claude Agent in IDEs on the Agent SDK | dev tools | Agent SDK, Claude Code, MCP | IDEs trusted by 15+ million developers worldwide | | | [Jumpcut](https://claude.com/customers/jumpcut) | Script coverage generated in seconds | media | API | 35,000+ hours saved in script reading time | | | [Juno](https://claude.com/customers/juno) | Voice and text agents for daily patient care | healthcare | API, Claude Code, Cowork | Onboarded more than 100,000 patients since launching in October | ✓ | | [Kai](https://claude.com/customers/kai) | Autonomous triage of millions of security findings | security | API, Bedrock | 99.5% of 2.5 million software composition analysis findings eliminated as false positives | | | [Kindora](https://claude.com/customers/kindora) | Non-engineer founder built a grant-matching SaaS | nonprofit | Claude Code, API, MCP | 328 nonprofits on the platform within months of beta launch | | | [Kodif](https://claude.com/customers/kodif) | Autonomous support agents that handle refunds | customer support | API, Bedrock | 90% automation for Trust Wallet crypto support | | | [L'Oréal](https://claude.com/customers/loreal) | Multi-agent conversational analytics for employees | retail | API, Claude Code | 44,000 monthly users across L'Oréal's AI platform with 2.5 million messages per month | | | [LaunchNotes / Graph](https://claude.com/customers/graph) | Sprint analyses and release notes from dev data | dev tools | API, Bedrock | 5x faster incident identification for SRE teams | | | [Law&Company](https://claude.com/customers/law-and-company) | Korean legal drafting and research in SuperLawyer | legal | API | Attracted 6,000 users, 20% of South Korean Practicing lawyers, within 180 days | | | [Lazy AI](https://claude.com/customers/lazy-ai) | Full-stack app generation from conversation | dev tools | API | 51% reduction in code requiring multiple fixes | | | [Legora](https://claude.com/customers/legora) | Assistant, tabular review, and legal workflows | legal | API, Claude Code | 18% higher performance on large legal evaluation set | ✓ | | [Lex](https://claude.com/customers/lex) | In-document AI feedback for writers | media | API | 25,000 signups within 24 hours of launch | | | [LG CNS](https://claude.com/customers/lg-cns) | Software factory migrating legacy enterprise systems | consulting | Claude Code, Bedrock | Converted 2,888 of 2,913 APIs from a 20-year-old legacy system, a 99.1% completion rate | | | [Lindy](https://claude.com/customers/lindy) | Agents for outreach, support, and scheduling | sales & marketing | API | Reduce time-to-qualified-lead by 72% for sales teams | | | [Local Falcon](https://claude.com/customers/local-falcon) | Bulk review analysis and local SEO recommendations | sales & marketing | API | 1 million customer reviews analyzed simultaneously | | | [Lokalise](https://claude.com/customers/lokalise) | AI translation suggestions inside localization workflows | other | API | 82.6% acceptance rate for AI translation suggestions with Claude | | | [Lotte Homeshopping](https://claude.com/customers/lotte-homeshopping) | Moni answers partner compliance questions 24/7 | retail | API | Reduced partner inquiries to QA staff by 30-40% | | | [Lovable](https://claude.com/customers/lovable) | Agentic app builder from conversation | dev tools | API | $200M ARR in a year after launch | | | [Lyft](https://claude.com/customers/lyft) | Support assistant resolving rider issues in seconds | other | API | Reduced customer support resolution time by over 87% | | | [MagicSchool](https://claude.com/customers/magicschool) | Teacher artifact generation and student feedback | education | API, Bedrock, Vertex | 7 million educators actively using platform | | | [MagicSchool (Q&A)](https://claude.com/customers/magicschool-qa) | Real-time student message safety moderation | education | API, Claude Code | 8 to 10 million student messages moderated in real time each month | | | [Matillion](https://claude.com/customers/matillion) | Maia builds and migrates data pipelines | dev tools | Claude Code | Sophisticated data transformations that took a customer 40 hours now take 1 hour with Maia | | | [Medgate](https://claude.com/customers/medgate) | Spec-to-feature development for telemedicine | healthcare | Claude Code | Up to 80% faster test case development | | | [micro1](https://claude.com/customers/micro1) | AI technical interviews at scale | HR & recruiting | API | 3,000+ interviews conducted daily by AI | | | [Mintlify](https://claude.com/customers/mintlify) | Docs Q&A assistant and drift-fixing agent | dev tools | API, Claude Code | 67% of documentation queries resolved by AI assistant without human intervention | ✓ | | [Money Forward](https://claude.com/customers/money-forward) | Claude Code as the daily engineering driver | fintech | Claude Code, MCP | Reduced API endpoint implementation by 70%, from 2 days to 5 hours | | | [Mutiny](https://claude.com/customers/mutiny) | Three agents one-shot branded sales assets | sales & marketing | API | 3x improvement in design satisfaction since making Claude Opus the default model | ✓ | ## N to Z | Company | What they built | Vertical | Surface | Standout number | YC | |---|---|---|---|---|---| | [N26](https://claude.com/customers/n26) | Chargeback analysis and document-heavy banking workflows | fintech | API, Bedrock, Agent SDK | Automated up to 70% of tasks across targeted processes, with ongoing improvements | | | [n8n](https://claude.com/customers/n8n) | AI Workflow Builder assembling automations in chat | productivity | API, Claude Code | Cuts 80% of the work out of building a workflow | | | [National Domestic Workers Alliance](https://claude.com/customers/national-domestic-workers-alliance) | Ask Aya guides domestic workers in real time | nonprofit | API | 93% of beta testers took action on Ask Aya's advice | | | [National Domestic Workers Alliance (Q&A)](https://claude.com/customers/national-domestic-workers-alliance-qa) | Worker-governed AI self-advocacy support | nonprofit | API | 93% of beta testers acted on Ask Aya's advice and 25% negotiated increases in compensation | | | [Nevis (Nevis Wealth)](https://claude.com/customers/nevis) | Meeting agents and smart tasks for advisors | fintech | API, Claude Code | 5+ hours per week saved per advisor on administrative tasks | | | [Newfront](https://claude.com/customers/newfront) | Benefits assistant and contract review | fintech | API | 60% cost savings in document processing through automation | ✓ | | [Nomura Research Institute / NRI](https://claude.com/customers/nri) | Checklist review of Japanese business documents | consulting | API, Bedrock | 50% faster reviews for complex Japanese business documents | | | [Norges Bank Investment Management (NBIM)](https://claude.com/customers/nbim) | Enterprise-wide research synthesis and analysis | fintech | Enterprise, Claude Code, MCP | 20% time saved weekly per employee on Claude assisted analytical and operational tasks | | | [Notion](https://claude.com/customers/notion) | Managed Agents run tasks from Notion boards | productivity | Managed Agents | 90% cost reduction with prompt caching while cutting latency by up to 85% | | | [Notion (Q&A)](https://claude.com/customers/notion-qa) | Task board orchestrates long-running agent sessions | productivity | Managed Agents | 30+ concurrent agent tasks from a single task board | | | [Novo Nordisk](https://claude.com/customers/novo-nordisk) | NovoScribe generates regulatory clinical documents | healthcare | Claude Code, Bedrock | Time spent producing clinical study documentation reduced from 10+ weeks to 10 minutes | | | [OffDeal](https://claude.com/customers/offdeal) | Archie runs the M&A deal lifecycle | fintech | Agent SDK, API, MCP | Increased eval accuracy from 25% to 85% after switching to the Claude Agent SDK | ✓ | | [OpusClip](https://claude.com/customers/opusclip) | GTM staff build call-intelligence workflows | media | Claude Code, API, MCP | 100% automated sales call review coverage up from 5–10% manual review | | | [Orange](https://claude.com/customers/orange) | Manga localization applied onto page images | media | API | 5-7x efficiency improvements over existing processes | | | [Otter](https://claude.com/customers/otter) | Meeting summaries, action items, and chat | productivity | API, Bedrock, Vertex | 50 million meetings summarized annually | | | [Pacific Community Ventures](https://claude.com/customers/pacific-community-ventures) | Voice survey tagging for impact research | nonprofit | API | ~300 workers reached in one nationwide voice survey, up from 12 in a comparable focus group | | | [Pacific Community Ventures (Q&A)](https://claude.com/customers/pacific-community-ventures-qa) | Qualitative signals feed a community lending credit model | nonprofit | Bedrock, Vertex, API | Pacific Community Ventures scales worker feedback 10x with Claude | | | [Palo Alto Networks](https://claude.com/customers/palo-alto-networks) | In-IDE completion and CI/CD post-processing | security | API, Vertex | 20-30% increase in feature development velocity | | | [Panorama](https://claude.com/customers/panorama) | Student data summaries and intervention recommendations | education | API, Bedrock | 25% of U.S. student population served | | | [Panther](https://claude.com/customers/panther) | Alert triage agents inside customer AWS environments | security | API, Bedrock | 70% reduction in alert fatigue through automated triage | | | [Parcha](https://claude.com/customers/parcha) | Due diligence and transaction monitoring agents | fintech | API, Agent SDK | 3 months to 5 minutes Customer due diligence workflow time reduced | | | [Pelanor](https://claude.com/customers/pelanor) | Cloud cost anomaly explanations tied to code changes | dev tools | API, Claude Code | 10x growth from 5 to 50 customers in 6 months | | | [Pendo](https://claude.com/customers/pendo-qa) | Novus detects usability issues and generates fix PRs | dev tools | Managed Agents, Claude Code, MCP | 90% success rate On PM-reviewed evaluation sets for Claude-powered agent tasks | | | [Pensive](https://claude.com/customers/pensive) | Rubric-based exam grading and a 24/7 tutor | education | API | 7% increase in student midterm scores in large intro CS courses | | | [Perplexity](https://claude.com/customers/perplexity) | Cited answers across search tiers | productivity | API, Bedrock | 2x faster response times with Claude 3.5 Sonnet | | | [Postman](https://claude.com/customers/postman) | Agent Mode generates collections, tests, and specs | dev tools | API, Bedrock, Claude Code | Up to 1,150 hours saved per year for developers using Agent Mode on average | | | [Pratham International](https://claude.com/customers/pratham-international) | Generates, digitizes, and grades student assessments | education | API | 1,500+ student assessments completed across 20 schools | | | [Praxis AI](https://claude.com/customers/praxis) | Professor digital-twin teaching assistants | education | API, Bedrock | 75% student engagement with digital twins, compared to 14% for generic AI tools | | | [Presien](https://claude.com/customers/presien) | /loop reasons over live worksite safety data | other | API, MCP | Over 70% reduction in critical safety events within the first three months of deployment | | | [Pressmaster.ai](https://claude.com/customers/pressmaster) | Personal Twin generates thought-leadership content | sales & marketing | API, Claude Code | 90%+ reduction in content production time | | | [Pulpit AI](https://claude.com/customers/pulpit-ai) | One sermon becomes 20+ content pieces | media | API | 200% increase in customer base within 3 months | | | [PwC](https://claude.com/customers/pwc-qa) | Legacy code analysis and client deliverables | consulting | Claude Code, Enterprise | What used to be six weeks' worth of engagements is now about two and a half days | | | [Qodo](https://claude.com/customers/qodo) | PR review, test generation, and codebase comprehension | dev tools | API | 1 million pull requests reviewed per quarter | | | [Qualified Health](https://claude.com/customers/qualified-health) | Population-scale screening for eligible patients | healthcare | API | [see note](/case-studies/strength-of-evidence) | | | [Quantium](https://claude.com/customers/quantium) | Company-wide assistant for code and proposals | consulting | API, Enterprise | 89% of team use AI daily in their work | | | [Quantium](https://claude.com/customers/quantium-qa) | Enterprise-wide Claude plus client agent builds | consulting | Enterprise, Claude Code, Cowork | 1,200+ Claude Enterprise users across Quantium | | | [Quillit](https://claude.com/customers/quillit) | Interview transcript summarization with citations | consulting | API | 80% reduction in report writing time | | | [RAINN](https://claude.com/customers/rainn) | Encrypted survivor-contact channel integrations | nonprofit | Claude Code, Enterprise, Bedrock | With Claude Code, we shipped the full Signal integration in roughly 30 days. | | | [Rakuten](https://claude.com/customers/rakuten) | Autonomous coding plus specialist Managed Agents | retail | Claude Code, Managed Agents | 79% reduction in time to market (from 24 days to 5 days) | | | [Rakuten](https://claude.com/customers/rakuten-qa) | Cloud-hosted specialist agents beside employees | retail | Managed Agents, API, Claude Code | 97% reduction in initial critical errors | | | [Ramp](https://claude.com/customers/ramp) | Agentic coding from tickets to incident response | fintech | Claude Code, MCP | 1M+ lines of AI code implemented in 30 days | | | [Replit](https://claude.com/customers/replit) | Agent 4 owns the full development lifecycle | dev tools | API, Vertex | 6+ hours continuous autonomous development without human intervention | ✓ | | [Reversia](https://claude.com/customers/reversia) | Whole-store Shopify translation into 110+ languages | retail | API | 99% translation accuracy validated by native-speaking translation professionals | | | [RileyBot](https://claude.com/customers/rileybot) | Safety-guardrailed K-12 tutor with parent logs | education | API | qualitative | | | [Rising Academies](https://claude.com/customers/rising-academies) | WhatsApp math tutoring across Africa | education | API | Achieved a 0.3 standard deviation effect size in learning outcomes | | | [Rocket](https://claude.com/customers/rocket) | One-shot multi-page website generation | dev tools | API, Agent SDK | 1M+ users building websites on the platform | | | [Rogo](https://claude.com/customers/rogo) | Research plus deck and model generation for bankers | fintech | API | 50,000+ queries per day across research, analysis, and artifact generation | | | [Satispay](https://claude.com/customers/satispay) | Claude Code across a mature payments codebase | fintech | Claude Code, Enterprise, Cowork | 75%+ of code committed each month is generated with Claude | | | [Scribd](https://claude.com/customers/scribd) | Bulk metadata generation for 100M+ documents | media | API | 70% of content improved with AI-generated metadata | | | [Section](https://claude.com/customers/section) | Internal strategy partner and the ProfAI coach | education | API | 82% of team members use Claude, compared to a 5% benchmark for the overall workforce | | | [Semgrep](https://claude.com/customers/semgrep) | False-positive filtering and one-click autofixes | security | API, Bedrock, MCP | Confidently labels 20% of security findings as safe to ignore, with a 92% user agree rate | | | [Sentry](https://claude.com/customers/sentry) | Seer root-causes errors and opens fix PRs | dev tools | Managed Agents, Vertex, Claude Code | Over 1 million RCAs (Root Cause Analysis) efficiently processed a year | | | [Sett](https://claude.com/customers/sett) | Orchestrator agent produces playable ads | sales & marketing | API, Bedrock | Reduced playable ad production time from one week to 75 minutes | | | [Shopify](https://claude.com/customers/shopify) | Sidekick turns merchant questions into analytics | retail | Claude Code, Vertex | qualitative | | | [Shortcut / Fundamental Research Labs](https://claude.com/customers/shortcut) | Multi-agent spreadsheet analysis and auditing | productivity | API, Claude Code | Benchmark accuracy from 7.29 to 8.08 out of 10 after swapping to Opus 4.6 with no prompt changes | | | [SK Telecom](https://claude.com/customers/skt) | In-call agent assist and post-call processing | other | API, Bedrock | 34% increase in LLM response quality ratings | | | [Skillfully](https://claude.com/customers/skillfully) | Job simulations that score demonstrated skills | HR & recruiting | API | 10x more likely to convert to full-time hires vs traditional methods | | | [Slack](https://claude.com/customers/slack) | AI search, summaries, and recaps in Slack | productivity | API, Claude Code | 97 minutes per week Saved by the average user through summarization and recap features | | | [Smartsheet](https://claude.com/customers/smartsheet) | MCP connector, Claude Code, and Enterprise rollout | productivity | API, Claude Code, Enterprise | 500+ engineers using Claude Code ship 3x more code and merge 31% more pull requests than peers on the same teams | | | [Snowflake](https://claude.com/customers/snowflake) | Cortex Analyst text-to-SQL over governed data | dev tools | API | 90% accuracy on complex text-to-SQL tasks | | | [Solvely.ai](https://claude.com/customers/solvely) | Step-by-step AI learning companions | education | API | #1 education app on App Store with 4.8/5 star rating | | | [Sourcegraph](https://claude.com/customers/sourcegraph) | Cody's default chat and completion model | dev tools | API | 75% increase in code insert rate | | | [Sourcegraph (Claude for Work)](https://claude.com/customers/sourcegraph-claude-for-work) | Community feedback synthesis for the product team | dev tools | Enterprise | 95% feedback accuracy in identifying board meeting issues | | | [Spotify](https://claude.com/customers/spotify) | Background agent runs fleet-wide code migrations | media | Agent SDK, API, Claude Code | Merged 650+ agent-generated pull requests into production per month | | | [Spring.new](https://claude.com/customers/spring-new) | Business apps built from natural-language prompts | dev tools | API, Vertex | 95-99% time savings on R&D projects | | | [Stairwell](https://claude.com/customers/stairwell) | Malware report summarization for analysts | security | API | 40,000+ characters processed in security data | | | [Steno](https://claude.com/customers/steno) | Transcript Genius searches deposition repositories | legal | API | 1200+ law firms using Steno services | | | [Stripe](https://claude.com/customers/stripe) | Claude Code pre-installed for every engineer | fintech | Claude Code, Agent SDK | Migrated 10,000 lines of Scala to Java in four days, a project estimated at ten engineering weeks | | | [StubHub](https://claude.com/customers/stubhub) | Virtual assistant resolves support end to end | media | API, Claude Code | Customer wait times dropped from over 20 minutes to near-instant responses | | | [StudyFetch](https://claude.com/customers/studyfetch) | Spark.E tutor generates study materials | education | API | 240 day streak maintained by students on learning platform | | | [Super Teacher](https://claude.com/customers/super-teacher) | First-draft educational games and lessons | education | API | 2x more productive engineering and content teams with Claude | | | [Syracuse University](https://claude.com/customers/syracuse) | University-wide Claude plus institutional data agents | education | Enterprise, Claude Code, MCP | 394% growth in student daily active users and 214% growth in staff month over month | | | [Syracuse University (Q&A)](https://claude.com/customers/syracuse-university) | Pedagogy redesign and Clementine class search | education | Claude Code, MCP | Exam scores were 12 points higher than they've ever been. | | | [Tabnine](https://claude.com/customers/tabnine) | Coding assistant for regulated customers | dev tools | API, Bedrock | Saw a 20% increase in free-to-paid user conversions | | | [Tahoe Lead Removal Project](https://claude.com/customers/tahoe-lead-removal-project) | Technical and regulatory analysis for cable removal | nonprofit | Enterprise | 6 miles of toxic lead cable removed from Lake Tahoe | | | [Tasklet](https://claude.com/customers/tasklet) | Long-running unattended business automations | productivity | API, MCP, Claude Code | 450,000 agent actions executed per day across all customer agents | ✓ | | [TELUS](https://claude.com/customers/telus) | Fuel iX lets 57,000 employees build AI solutions | other | API, Claude Code, MCP | 57,000 team members actively using generative AI | | | [The Epilepsy Foundation](https://claude.com/customers/epilepsy-foundation) | Sage answers grounded in approved epilepsy content | healthcare | API, Bedrock | 60,000 interactions with Sage reached in under a year, before any major promotion | | | [The Epilepsy Foundation (Q&A)](https://claude.com/customers/epilepsy-foundation-qa) | Grant writing and conversation analysis | healthcare | API | Without Claude, I would have likely had to spend 4x as much time on this grant. | | | [The Patrick J. McGovern Foundation](https://claude.com/customers/pjmf) | Grant Guardian standardizes nonprofit financials | nonprofit | Enterprise, Bedrock, Vertex | 100+ philanthropies using Grant Guardian across US | | | [Thomson Reuters](https://claude.com/customers/thomson-reuters) | CoCounsel legal and tax analysis with RAG | legal | API, Bedrock | 3,000+ subject matter experts' knowledge delivered | | | [Thomson Reuters (Cowork Q&A)](https://claude.com/customers/thomson-reuters-qa) | Cowork pilots for non-developer teams | legal | Cowork, Enterprise, Claude Code | qualitative | | | [Tidio](https://claude.com/customers/tidio) | Lyro resolves support chats end to end | customer support | API, Vertex | 71% automation of own customer support | | | [Tines](https://claude.com/customers/tines) | Natural-language automation for security analysts | security | API, Bedrock, MCP | 100x faster time-to-value, converting 120-step workflows to single-step agents | | | [tl;dv](https://claude.com/customers/tldv) | Cross-meeting insights and playbook scoring | sales & marketing | API | 300% growth in new customer sign-ups | | | [Tome](https://claude.com/customers/tome) | Account research and sales positioning briefs | sales & marketing | API | qualitative | | | [Trellix](https://claude.com/customers/trellix) | Wise triages security events autonomously | security | API, Bedrock | 8 hours saved for every 100 security alerts processed | | | [Triple Whale](https://claude.com/customers/triple-whale) | Analytics agent swarm over brand data | retail | API | 50%+ KPI increase in north star metrics for early customers | | | [TRY](https://claude.com/customers/try) | Enterprise Claude as an agency sparring partner | sales & marketing | Enterprise | 30% reduction in time spent on routine tasks | | | [Twilio](https://claude.com/customers/twilio-qa) | Solo PM built a spec-to-code Feature Factory | dev tools | Claude Code, MCP, Managed Agents | 425 API tools packaged into a Claude Code plugin | | | [Vambe](https://claude.com/customers/vambe) | WhatsApp sales agents that qualify and collect | sales & marketing | API, Claude Code | 95%+ multi-agent reliability, up from 30%, enabling multi-agent workflows | | | [Vanta](https://claude.com/customers/vanta) | Generates tailored compliance remediation steps | security | API, MCP | 113% increase in developer AI tool adoption in 2 months | ✓ | | [Vapi](https://claude.com/customers/vapi) | Composer builds voice agents conversationally | dev tools | API, Agent SDK, Claude Code | Made users nearly 3x more likely to reach 100 call minutes | ✓ | | [Vercel](https://claude.com/customers/vercel-qa) | Internal skill-powered agents and skills.sh | dev tools | Enterprise, Claude Code, MCP | ~100 skills power Vercel's internal data science agent | | | [Vibecode](https://claude.com/customers/vibecode) | Plain-language mobile apps published to the App Store | dev tools | API, Claude Code, Managed Agents | Reduced average app development cost from $50,000 to $100 | | | [Warp](https://claude.com/customers/warp) | Terminal agent defaults to Claude models | dev tools | API, Claude Code | 10M Claude Code sessions run inside Warp's terminal to date, including 400K+ each week | | | [Wedia Group](https://claude.com/customers/wedia-group) | Vision metadata for images at DAM scale | sales & marketing | API, Bedrock | 90% time reduction in metadata tagging with Claude | | | [Windsurf](https://claude.com/customers/windsurf) | Cascade in-IDE coding agent reasoning layer | dev tools | API, Bedrock, Vertex | 100M tokens processed per minute | | | [Wiz](https://claude.com/customers/wiz) | Claude Code migrated a 50,000-line library to Go | security | Claude Code, Bedrock, MCP | 50,000 lines of Python to Go in ~20 hours | | | [Wordsmith](https://claude.com/customers/wordsmith) | Contract bundles reviewed against compliance playbooks | legal | Claude Code, Bedrock | Processes 400-page contract bundles with 300-point compliance checks in 4-5 minutes | | | [Workato](https://claude.com/customers/workato) | Managed MCP servers for enterprise actions | productivity | API, MCP, Claude Code | 7x increase in employee adoption after deploying MCP servers | | | [WRTN](https://claude.com/customers/wrtn) | Interactive story characters and companions | media | API | 4.5 million monthly active users | | | [YMCA South Australia](https://claude.com/customers/ymca-south-australia) | Cross-site reports, tenders, and 20+ custom skills | nonprofit | Enterprise, MCP | 10-15 hours saved per week for key users | | | [Yoodli](https://claude.com/customers/yoodli) | AI roleplay personas for sales practice | sales & marketing | API | 23% more deals closed by reps who practice three or more roleplay scenarios per week | | | [You.com](https://claude.com/customers/you-dot-com) | Research Agent and enterprise workflow agents | productivity | API | 1000% revenue increase over the last year | | | [Zapia / BrainLogic](https://claude.com/customers/zapia) | WhatsApp shopping assistant for Latin America | retail | API, Vertex | 2.5 million users in first year across Latin America | | | [Zapier](https://claude.com/customers/zapier) | Company-wide Claude with 800+ internal agents | productivity | Enterprise, Claude Code, MCP | 89% AI adoption across all employees | | | [Zapier (Cowork)](https://claude.com/customers/zapier-cowork-qa) | Cowork runs cross-system knowledge work | productivity | Cowork, Enterprise, MCP | 15 SQL queries synthesizing live data from 6 engineering systems in one Cowork session | | | [Zencoder](https://claude.com/customers/zencoder) | Agent SDK core for enterprise coding agents | dev tools | Agent SDK, Claude Code, MCP | 2x reduction in AI code churn compared to Zencoder's previous best solution | | | [Zingage](https://claude.com/customers/zingage) | Voice agents restaff home-care shifts | healthcare | API | 82% reduction in after-hours labor costs at one Medicaid agency | | | [Zoom](https://claude.com/customers/zoom) | AI Companion summaries, Q&A, and docs | dev tools | API | 14% improvement in meeting summary accuracy | | ## Further Reading - [Case Studies overview](/case-studies) how the corpus is cut - [The startup cut](/case-studies/startups) companies that were small when they started - [Strength of evidence](/case-studies/strength-of-evidence) which numbers hold up, and which to avoid - [Reference Index](/reference) every published source, dated --- # Changelog URL: https://buildonanthropic.com/changelog import Changelog from '@site/src/components/Changelog'; # Changelog A log of every change to this wiki, newest first, grouped by month. Each page shows up here when it is added (**New**), every time it is edited (**Updated**), and when it is deleted (**Removed**). Each row shows the date, the section, the title, and the one-line description from frontmatter. Dates and changes are derived from git history (renames followed). Changes that have not been committed yet will not appear here. This is the full version of the "recent changes" list on the home page. --- # Agent Once, Session Every Run URL: https://buildonanthropic.com/concepts/agent-once-session-every-run # Agent Once, Session Every Run *The Managed Agents usage rule: create the agent, a versioned config, once; create a session for every run. Creating agents in the request path is the anti-pattern.* ![Hand-drawn ink-line illustration on terracotta: a loopy hand holds one large cream block aloft while a receding row of small identical cream blocks marches away below, a one-line profile face observing.](/img/pages/concepts-agent-once-session-every-run.jpg) --- ## The object model Managed Agents splits an agent into two objects, per the [platform docs](https://platform.claude.com/docs) (verify against platform.claude.com/docs; beta surfaces drift): - An **agent** is a persisted, versioned configuration: the model, the system prompt, the tools, the MCP servers, the skills. All of that lives on the agent, never on the session. Every update produces an immutable version. - A **session** is one run. It references an agent ID and an environment, streams events, and ends. The rule follows from the split: create the agent once, store its ID, and reuse it forever. Each incoming job creates only a session. ## Why request-path agent creation is the anti-pattern Calling agent creation inside the request path is the documented mistake, and it fails three ways at once. It orphans agents, one per request, until the account is a junkyard of unversioned configs. It spends creation latency on every run. And it defeats the versioning that is the point of the object model: with a stable agent ID, every session is traceable to an exact immutable config version, sessions can pin to a version, and a bad prompt change can be rolled back by pointing back at the previous one. The shape the docs recommend is a control plane and a data plane: the agent config lives in version-controlled YAML applied with the `ant` CLI, and application code touches only session creation. ## Why it earns a concept page This is the platform's version of a rule every infrastructure engineer already knows: configuration is not a runtime object. Teams coming from raw Messages API calls, where every request is self-contained, tend to carry that habit into Managed Agents and rebuild the agent per request because nothing stops them. The object model is the guardrail only if you use it as one. It also compounds with the rest of the platform. [Scheduled Deployments](/concepts/scheduled-deployments) fire sessions against a stable agent ID on a clock. [Memory Stores](/concepts/memory-stores) give those sessions continuity. And [prompt caching](/concepts/prompt-caching) rewards a frozen, versioned system prompt, since any byte of churn in the prefix invalidates the cache behind it. ## Who runs this shape in production - Rakuten deploys stable specialist agents and reports "We deploy each specialist agent within a week, managing long-running tasks across engineering, product, sales, marketing, and finance." ([source](https://claude.com/customers/rakuten-qa)) - Notion's task board fires sessions against its agents at volume: "30+ concurrent agent tasks from a single task board." ([source](https://claude.com/customers/notion)) ## Further Reading - [The Four Ways to Build](/concepts/four-ways-to-build): where Managed Agents sits among the paths. - [Scheduled Deployments](/concepts/scheduled-deployments): sessions fired on a clock against a stable agent. - [Memory Stores](/concepts/memory-stores): persistence across those sessions. - [Prompt Caching](/concepts/prompt-caching): why a frozen config is also a cheaper one. --- # Consumption Pricing URL: https://buildonanthropic.com/concepts/consumption-pricing # Consumption Pricing *The Claude Developer Platform bills by tokens consumed, not seats sold. The bill scales with how much the product actually runs.* ![Hand-drawn line illustration in warm terracotta and cream: two loopy line-drawn hands drawn from one continuous line, the left holding a small cream cup and the right a much larger one, a one-line profile face above between them.](/img/pages/concepts-consumption-pricing.jpg) --- ## The two meters Anthropic sells the same models through two different meters. Claude, Claude Code, and Claude for Work subscriptions are seats: a person pays a flat rate and usage lives inside plan limits. The API is consumption: per-token prices by model, published on the [platform pricing docs](https://platform.claude.com/docs) (verify current figures there; prices change with model generations), with no seat in sight. The distinction matters most to a founder because their product sits on the consumption meter. A seat business's cost is headcount-shaped and predictable; a token business's cost is usage-shaped and grows exactly as fast as the product succeeds. Every scaling story in the customer corpus is implicitly a consumption story: Tasklet reports "450,000 agent actions executed per day across all customer agents" ([source](https://claude.com/customers/tasklet)), and Windsurf reports "100M tokens processed per minute" ([source](https://claude.com/customers/windsurf)). Numbers like those are simultaneously growth metrics and cost structures. ## Why the alignment is real Under consumption pricing, the platform earns nothing extra from a builder's idle integration. Revenue arrives only when the builder's product runs, which means the platform's growth depends on builders' products being worth running at volume. That is a structurally different relationship than a seat license, where the sale precedes the usage and shelfware still bills. It also explains why the cost levers exist and are documented so aggressively. [Prompt caching](/concepts/prompt-caching) reads at a fraction of base price, [the Batch API](/concepts/the-batch-api) at half, [the effort dial](/concepts/the-effort-dial) trimming tokens per step: each one lowers the price of a unit of work, which lowers the threshold at which automating that work is worth it, which expands what gets built. Dust's reported "$10k/day saved" on model spend after optimizing caching ([source](https://claude.com/customers/dust)) is a customer keeping margin, and a workload that stays economical to run. ## The unit that matters is the outcome Token prices are the input cost; the number a business actually manages is cost per completed unit of work. The corpus's most instructive figures are stated that way: Freedom Forever prices its permit agent at "$0.20–$0.35" per run against a human-equivalent task it says costs several times more ([source](https://claude.com/customers/freedom-forever)), and OffDeal reports a buyer list that "would traditionally take a team of analysts more than a week to assemble and cost upward of $12,000" now produced "for roughly $200 in compute." ([source](https://claude.com/customers/offdeal)) When the per-outcome math clears, consumption growth is good news on both sides of the meter. When it does not, no discount rescues it; the fix is architecture, which is what the rest of the cost concepts are for. ## Further Reading - [Prompt Caching](/concepts/prompt-caching), [The Batch API](/concepts/the-batch-api), and [The Effort Dial and Model Splitting](/concepts/the-effort-dial): the three levers on the consumption bill. - [The Startup Cut](/case-studies/startups): small teams whose unit economics are published. - [Strength of Evidence](/case-studies/strength-of-evidence): how to read published cost numbers responsibly. --- # The Four Ways to Build URL: https://buildonanthropic.com/concepts/four-ways-to-build # The Four Ways to Build *The four paths onto the Claude Developer Platform, separated by two questions: who supplies the harness, and who supplies the deployment.* ![Hand-drawn ink-line illustration on terracotta: a loopy hand places the last of four cream squares into a two-by-two arrangement, a one-line profile face observing.](/img/pages/concepts-four-ways-to-build.jpg) --- ## The 2x2 Anthropic's platform offers four distinct ways to build an agent. They stop being confusing the moment you sort them by [the harness](/concepts/the-harness) (the agent loop) and [the deployment](/concepts/the-deployment) (the infrastructure it runs on). Details below are drawn from [platform.claude.com/docs](https://platform.claude.com/docs) and [code.claude.com/docs](https://code.claude.com/docs/en/agent-sdk); verify against the live docs, beta surfaces drift. | Path | You write | Harness | Deployment | |---|---|---|---| | **Messages API, manual loop** | the tool-use loop itself | you | you | | **Messages API + Tool Runner** | just your tool functions | the SDK | you | | **Claude Agent SDK** | a prompt plus options | the SDK (the Claude Code harness as a library) | you | | **Managed Agents** (beta) | an agent config | Anthropic | Anthropic | Three of the four leave deployment to the builder. Managed Agents is the only path where Anthropic supplies both halves: hosted sandbox, durable sessions, [Vaults](/concepts/vaults) for secrets, [Scheduled Deployments](/concepts/scheduled-deployments), capacity. One confusion worth killing early: **Tool Runner is not the Agent SDK.** Tool Runner is a helper inside the regular Anthropic SDK that loops over tools you define. The Agent SDK is Claude Code packaged as a library, with built-in file, shell, and web tools. Both are harness-only. ## Choosing between them - **One well-scoped call** (classify, extract, summarize): just the Messages API. Do not build an agent. This is Anthropic's own standing advice; [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) (December 2024) urges "finding the simplest solution possible, and only increasing complexity when needed." - **Custom tools, API-first, you want the loop:** Messages API with Tool Runner. - **Fast local iteration, or production on infrastructure you want to control:** Agent SDK, plus your own sandbox, session store, and isolation. - **Production without the infrastructure lift:** Managed Agents. ## What switching paths has done in practice - OffDeal moved its M&A agent to the Agent SDK and reports eval accuracy rising from 25% to 85% overall, with the switch itself doing the first leg: "Just switching to the SDK moved our score on that eval from 25% to 60%." ([source](https://claude.com/customers/offdeal)) - Freedom Forever benchmarked the Agent SDK on web automation at a "97.5% success rate on complex web automation benchmarks where alternative frameworks hit 20–60%." ([source](https://claude.com/customers/freedom-forever)) - Pendo tried building its own agent infrastructure for about three months, then reached production on Managed Agents in "About three days." ([source](https://claude.com/customers/pendo-qa)) - Sentry shipped its Managed Agents integration "in weeks instead of months" with one engineer. ([source](https://claude.com/customers/sentry)) ## The honest caveats Managed Agents is beta and, at time of writing, not available through Bedrock or Vertex; confirm current status in the [platform docs](https://platform.claude.com/docs) before committing an architecture to it. And sometimes the right tier is the lowest one: if a script or a single call solves the problem, every row above it is overhead. The wiki's [Build on Claude tool](/tools) walks a specific product through this exact decision. ## Further Reading - [The Harness](/concepts/the-harness) and [The Deployment](/concepts/the-deployment): the two axes of the 2x2. - [Agent Once, Session Every Run](/concepts/agent-once-session-every-run): the usage rule once you land on Managed Agents. - [The Tools](/tools): a paste-in file that applies this decision to your product. - [The Index](/case-studies/the-index): every published build, with the surface each one chose. --- # Concepts URL: https://buildonanthropic.com/concepts # Concepts *The lexicon. Nineteen concepts in four groups: the structural spine, the Managed Agents primitives, the cost levers, and the patterns proven in 256 published customer stories.* ![Hand-drawn line illustration in warm terracotta and cream: two loopy hands sort small cream tiles into neat square clusters, observed by a face drawn in a single line.](/img/pages/concepts-index.jpg) --- Each concept is one short page: a quotable definition line, a body that earns the term with named sources, and Further Reading that makes the lexicon a graph. Case-study claims quote the company's own published page and link to it; API details defer to [platform.claude.com/docs](https://platform.claude.com/docs), which is always the authority. ## The structural spine The vocabulary for the platform's one big architectural question: what do you own, and what do you delegate? - [The Harness](/concepts/the-harness): *the software loop around the model: tool dispatch, context management, retries, and compaction. The half of an agent that is code rather than model.* - [The Deployment](/concepts/the-deployment): *everything an agent runs on: sandboxing, session persistence, tenant isolation, secrets handling, and capacity.* - [The Four Ways to Build](/concepts/four-ways-to-build): *manual loop, Tool Runner, Agent SDK, Managed Agents, separated by who supplies the harness and who supplies the deployment.* - [Agent Once, Session Every Run](/concepts/agent-once-session-every-run): *create the versioned agent config once, create a session per run; agent creation in the request path is the anti-pattern.* ## The Managed Agents primitives What the managed platform runs so a small team does not have to build it. - [Vaults](/concepts/vaults): *secrets never enter the sandbox; they are substituted into outbound requests at egress, so a prompt injection has nothing to steal.* - [Outcomes](/concepts/outcomes): *define done as a rubric, and the platform runs an iterate, grade, revise loop with a separate grader until the work passes.* - [Memory Stores](/concepts/memory-stores): *workspace-scoped persistent memory mounted into every session, versioned and auditable.* - [Scheduled Deployments](/concepts/scheduled-deployments): *cron for agents: a schedule fires a fresh session on the clock, and every firing leaves a run record.* - [The Human Gate](/concepts/the-human-gate): *an approval step the platform enforces before an action commits; a prompt can ask for approval, a permission policy guarantees it.* ## Cost and economics The levers that decide whether an agent product has margins. - [Prompt Caching](/concepts/prompt-caching): *cached input reads cost a fraction of base price, but the cache is a strict prefix match: one changed byte invalidates everything after it.* - [The Batch API](/concepts/the-batch-api): *50 percent off input and output tokens for anything that can wait up to 24 hours; stacks with caching.* - [The Effort Dial and Model Splitting](/concepts/the-effort-dial): *frontier-model tokens for judgment, small-model tokens for commodity steps, effort turned down where thinking is wasted.* - [Consumption Pricing](/concepts/consumption-pricing): *the platform bills by tokens consumed, not seats sold, so the bill scales with how much the product actually runs.* ## Patterns proven in the case studies Recurring shapes in the published customer corpus, each cited to the companies that shipped it. - [The Company Brain](/concepts/the-company-brain): *a persistent store of how a specific company works, which agents read on every run and write back to as they learn.* - [The No-API Integration](/concepts/the-no-api-integration): *operating the screens of systems that will never get an API, with computer use standing in for the integration nobody will ship.* - [Silent Failure](/concepts/silent-failure): *the failure mode that kills trust in unattended agents: the agent thinks it finished, and it did not.* - [Self-Correction Mid-Task](/concepts/self-correction): *an agent that checks its work against reality before handing it over, and fixes what it finds.* - [The Audit Trail](/concepts/the-audit-trail): *citation to source as a product feature, so the human verifies instead of hunts.* - [Long-Running Work](/concepts/long-running-work): *agent runs measured in hours rather than seconds: the workload class that makes every architecture decision matter.* ## How to use the lexicon - **As a glossary.** Look up a term, read the italic line, move on. The [Glossary](/reference/glossary) holds the flattened one-line-per-term view. - **As an architecture course.** Read the spine in order, then the primitives, then the cost levers. That sequence is the platform decision a founder actually faces. - **As an evidence trail.** The pattern pages quote real published customer stories; every number links to the official page, and [Strength of Evidence](/case-studies/strength-of-evidence) explains how hard each kind of number is. ## Further Reading - [The Index](/case-studies/the-index): the 256 published stories the pattern pages cite from. - [Strength of Evidence](/case-studies/strength-of-evidence): how to weigh the numbers quoted across this section. - [Reference Index](/reference): the dated master list of Anthropic's published sources. - [The Tools](/tools): the paste-in generator that applies these concepts to your product. --- # Long-Running Work URL: https://buildonanthropic.com/concepts/long-running-work # Long-Running Work *Agent runs measured in hours rather than seconds: the workload class that makes every harness and deployment decision matter.* ![Hand-drawn ink-line illustration on terracotta: a loopy hand draws one very long meandering line across the frame, passing a cream sun disc and a cream crescent moon, ending at a one-line profile face.](/img/pages/concepts-long-running-work.jpg) --- ## The workload class A chat completion lives for seconds; a long-running agent works for hours across hundreds of tool calls, unattended. The newest builder cohort treats this as a baseline requirement: in YC's Summer 2026 batch, 17 percent of companies describe long-running, scheduled, or background execution, with machine0 stating the shape plainly in its launch material: "a coding agent debugging a complex issue runs 4-8 hours," and Rex describing "agents run continuously in the background." ## The published runs The customer corpus has real numbers on how long production runs already go: - Rakuten reports "7 hours of sustained autonomous coding on a complex open-source refactoring project" in vLLM, a 12.5 million line codebase, with the result at "99.9% numerical accuracy" against the reference method. The engineer's account: "I didn't write any code during those seven hours. I just provided occasional guidance." ([source](https://claude.com/customers/rakuten)) - Replit reports "Sessions run for 6+ hours without human input, a 10x improvement over previous agent capabilities," and a power user who "generated more than 36,000 lines of production-ready code in a single Agent 4 session" over "roughly 400 minutes of autonomous runtime." ([source](https://claude.com/customers/replit)) - OffDeal "Built a buyer sourcing agent that runs autonomously for up to 4 hours, researching potential acquirers across 10 different sourcing methods." ([source](https://claude.com/customers/offdeal)) - Tasklet runs "long-running, unattended multi-step business automations in the cloud" as its entire product, at "450,000 agent actions executed per day across all customer agents." ([source](https://claude.com/customers/tasklet)) ## What hours-long runs demand Every concept in this lexicon gets stress-tested by run length. Anthropic's [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) (November 2025) supplies the harness-side toolkit: an initializer agent that sets up the environment, progress files the agent maintains as an external memory of where it is, and environment management so a dying process does not end the mission. Its context-engineering guidance adds compaction, since no window holds an eight-hour transcript, and warns about context rot, the quality decay of an ever-fuller window. [Self-correction](/concepts/self-correction) is what keeps hour six on course with hour one, and [silent failure](/concepts/silent-failure) is what an unwatched run risks the whole time. The deployment side is just as implicated: a session that runs for hours needs [durable state that survives a restart](/concepts/the-deployment), capacity held for the duration, and, if it runs on a clock, [a schedule with per-firing run records](/concepts/scheduled-deployments). ## Why it is the decisive workload Short tasks forgive weak architecture; long ones do not. This is the workload class where the [harness and deployment](/concepts/four-ways-to-build) choice stops being a preference and becomes the product, which is why the stories above cluster on the Agent SDK and Managed Agents rather than hand-rolled loops. If your roadmap ends at runs measured in hours, the architecture conversation starts here. ## Further Reading - [The Harness](/concepts/the-harness): the loop that has to survive the hours. - [The Deployment](/concepts/the-deployment): the state and capacity underneath. - [Memory Stores](/concepts/memory-stores) and [Scheduled Deployments](/concepts/scheduled-deployments): continuity across and between runs. - [Self-Correction Mid-Task](/concepts/self-correction): what makes run length usable. --- # Memory Stores URL: https://buildonanthropic.com/concepts/memory-stores # Memory Stores *Workspace-scoped persistent memory for Managed Agents: a filesystem mounted into every session, versioned and auditable, so what an agent learns survives the run.* ![Hand-drawn line illustration in warm terracotta and cream: a loopy line-drawn hand lifts one cream folder out of a neat stacked column of identical folders resting on a single wobbly shelf line, a one-line profile face looking on.](/img/pages/concepts-memory-stores.jpg) --- ## What it is A session is disposable by design; a Memory Store is what persists across them. Per the [Managed Agents docs](https://platform.claude.com/docs) (verify against platform.claude.com/docs; beta surfaces drift), a Memory Store is workspace-scoped persistent storage mounted into the session's container as a filesystem, at a path like `/mnt/memory//`. The agent reads and writes it with ordinary file tools. Contents are versioned, carry an audit trail, and support redaction. ## Why files, not a database Anthropic's context-engineering guidance treats structured note-taking as a first-class memory technique: [Effective context engineering for AI agents](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) (September 2025) frames context as "a finite resource with diminishing marginal returns," which means an agent cannot simply carry everything it knows in the window. The working alternative is notes on disk: the agent writes down what it learned, and the next session reads only what it needs. A mounted filesystem makes that pattern native. No retrieval service, no embedding pipeline, no schema migration; the model already knows how to use files. The versioning and audit trail matter because memory is state that compounds. A wrong fact written today poisons every future session that reads it, so you want to see what was written, when, and by which run, and be able to roll it back. ## What it enables Persistent memory is the mechanism behind [the Company Brain](/concepts/the-company-brain), the pattern where an agent accumulates knowledge of how a specific organization works. It is also load-bearing for [long-running work](/concepts/long-running-work) and anything on a schedule: a nightly job that cannot remember last night starts from zero forever. - Rakuten's Managed Agents deployment is built on exactly this: "Cloud-hosted specialist agents with persistent compute and memory that work alongside employees across engineering, product, sales, marketing, and finance." ([source](https://claude.com/customers/rakuten-qa)) - Notion pairs agent orchestration with a "Self-improving skills database" that is "maintained by Claude after every completed task," so lessons from each run feed the next one. ([source](https://claude.com/customers/notion-qa)) - Greptile's review agent uses sub-agents for memory retrieval so codebase knowledge accumulated across reviews stays available without flooding the main context. ([source](https://claude.com/customers/greptile)) ## The discipline it requires Memory is not free context. Everything above about context as a scarce resource still applies to what gets read back in; a memory store that grows without curation becomes a slower way to be confused. The pattern that works in Anthropic's guidance is deliberate: the agent maintains specific, named artifacts (a progress file, a decisions log, a map of the territory) rather than an append-only transcript of everything it ever thought. ## Further Reading - [The Company Brain](/concepts/the-company-brain): the product pattern built on persistent memory. - [Long-Running Work](/concepts/long-running-work): why continuity across runs is a workload requirement. - [Agent Once, Session Every Run](/concepts/agent-once-session-every-run): the session model this persists across. - [The Harness](/concepts/the-harness): where context management lives. --- # Outcomes URL: https://buildonanthropic.com/concepts/outcomes # Outcomes *The Managed Agents primitive for defining done: you supply a rubric, and the platform runs an iterate, grade, revise loop until the work passes, graded by a separate context.* ![Hand-drawn line illustration in warm terracotta and cream: a loopy line-drawn hand holds a cream card up against a taller cream tablet marked with three checkmarks, two more cream cards waiting below, a one-line profile face observing.](/img/pages/concepts-outcomes.jpg) --- ## What it is An Outcome is a definition of done attached to a session. Per the [Managed Agents docs](https://platform.claude.com/docs) (verify against platform.claude.com/docs; beta surfaces drift), you send the session an outcome with a rubric, and the harness runs a loop: the agent works, a grader scores the result against the rubric, and if it fails, the agent revises. The loop defaults to 3 iterations with a maximum of 20. The grader is separate from the worker and holds its own context. ## Why the separate grader is the point Anthropic's guidance keeps returning to one failure mode: the builder must not grade its own work, because a model reviewing its own output inside the same context tends toward self-preferential bias. The evaluator-optimizer pattern in [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) (December 2024) and the trajectory-grading guidance in [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) (January 2026) both split generation from judgment. Outcomes bakes that split into the platform: the grader arrives fresh, reads the rubric and the work, and has no stake in the draft it is scoring. This is also the platform's structural answer to [silent failure](/concepts/silent-failure). An agent that decides for itself that it finished is exactly the agent that thinks it finished and did not. An agent whose finish line is a rubric applied by someone else has to earn the word done. ## What it replaces Teams building on the raw API assemble this loop by hand: a generator call, a judge prompt, retry logic, iteration caps, logging of each pass. It is buildable, and the published evidence shows how much the loop matters when built well: - Pendo evaluates its Managed Agents-powered Novus against PM-reviewed standards and reports "The agent had a 90%+ success rate against that standard." ([source](https://claude.com/customers/pendo-qa)) - OffDeal credits eval-driven iteration, not the model alone, for its climb: "Increased eval accuracy from 25% to 85% after switching to the Claude Agent SDK and iterating on MCP design and prompting." ([source](https://claude.com/customers/offdeal)) - CircleCI's Chunk agent closes the loop against reality itself, validating each fix in a real CI pipeline before delivering a green pull request. ([source](https://claude.com/customers/circleci)) ## How to use it well The rubric is the product surface. A vague rubric ("the report should be good") grades nothing; a concrete one ("every claim carries a source link; totals reconcile to the input ledger") turns the loop into leverage. Write the rubric the way you would brief a contractor whose work you will not re-check, because that is the operational bet you are making. Anthropic's own [launch-your-agent](https://github.com/anthropics/launch-your-agent) reference skill treats this as the core move: it grades a founder's new agent "against your own definition of done." ## Further Reading - [Silent Failure](/concepts/silent-failure): the trust problem this primitive exists to close. - [Self-Correction Mid-Task](/concepts/self-correction): the same instinct applied inside a run. - [Agent Once, Session Every Run](/concepts/agent-once-session-every-run): where outcomes attach in the object model. - [Strength of Evidence](/case-studies/strength-of-evidence): why eval-backed numbers are the ones worth quoting. --- # Prompt Caching URL: https://buildonanthropic.com/concepts/prompt-caching # Prompt Caching *Claude bills cached input tokens at a fraction of base price, but the cache is a strict prefix match: one changed byte invalidates everything after it.* ![Hand-drawn ink-line illustration on terracotta: two loopy hands square a stack of cream sheets flush at the left edge while one sheet juts out of line, a one-line profile face observing.](/img/pages/concepts-prompt-caching.jpg) --- ## The one rule Anthropic's post [Lessons from building Claude Code: prompt caching is everything](https://claude.com/blog/lessons-from-building-claude-code-prompt-caching-is-everything) (April 2026) states it as bluntly as documentation gets: "Prompt caching is a prefix match. Any change anywhere in the prefix invalidates everything after it. Design your entire system around this constraint." The same post notes the team runs alerts on its own cache hit rate and declares SEVs when it drops. That is caching treated as architecture, not as an optimization pass. The design consequence is an ordering discipline: stable content first. The request renders as tools, then system, then messages, per the [prompt caching docs](https://claude.com/blog/prompt-caching), so tool definitions and the system prompt should be frozen, and anything volatile (timestamps, request IDs, per-user detail) belongs after the last cache breakpoint. A "helpful" dynamic date string at the top of a system prompt silently zeroes the cache on every request behind it. ## The economics Per the [platform docs](https://platform.claude.com/docs) (verify against platform.claude.com/docs; pricing and TTLs drift): cache reads cost roughly 0.1x base input price. Cache writes cost 1.25x base for the 5-minute TTL or 2x for the 1-hour TTL, which puts break-even at the second request on the short TTL. For agents, whose loops resend a growing conversation on every model call, this is the difference between linear and quadratic-feeling input bills. Verification is one field: `usage.cache_read_input_tokens` in the response. If it reads zero across requests that should share a prefix, something in the prefix is churning, and the bill is quietly paying full price. ## The published numbers Caching produces some of the hardest cost figures in the entire customer corpus: - Notion reports "Reduced costs by 90% and latency by up to 85% with prompt caching." ([source](https://claude.com/customers/notion)) - Dust reports "Cache reads doubled from about 30% to 65% of input tokens, input spend fell 22%, and overall model spend fell 18–19%: roughly $10K saved per day." ([source](https://claude.com/customers/dust)) - Greptile "Achieves ~90% cache hit rates, dramatically reducing costs for both Greptile and self-hosting customers." ([source](https://claude.com/customers/greptile)) - Bolt's maker StackBlitz "sees roughly 90% cache efficiency on the Agent SDK, keeping inference costs manageable across these long-running workflows." ([source](https://claude.com/customers/bolt)) Note what the strongest cost stories have in common: they talk about cache hit rates and prompt structure, not about switching to a cheaper model. ## Where it comes free On Managed Agents, caching on repeated history is applied by the platform, and the Agent SDK's harness caches aggressively by default (the ~90% figures from Greptile and Bolt above are SDK deployments). On the raw Messages API you place cache breakpoints yourself, which is more control and more ways to break it. Either way the prefix rule governs, which is one more argument for [a frozen, versioned agent config](/concepts/agent-once-session-every-run). ## Further Reading - [The Batch API](/concepts/the-batch-api): the other big discount, and it stacks with this one. - [The Effort Dial and Model Splitting](/concepts/the-effort-dial): the remaining cost levers. - [Consumption Pricing](/concepts/consumption-pricing): why these levers are the platform's own economics. - [The Harness](/concepts/the-harness): the layer that does the caching for you. --- # Scheduled Deployments URL: https://buildonanthropic.com/concepts/scheduled-deployments # Scheduled Deployments *Cron for agents: a schedule fires a fresh session on the clock, and every firing leaves a run record.* ![Hand-drawn line illustration in warm terracotta and cream: a large cream clock face with two plain hands and no numbers, a row of four evenly spaced cream squares marching beneath it, a loopy line-drawn hand resting beside the clock and a one-line profile face watching.](/img/pages/concepts-scheduled-deployments.jpg) --- ## What it is A Scheduled Deployment attaches a cron schedule to an agent. Per the [Managed Agents docs](https://platform.claude.com/docs) (verify against platform.claude.com/docs; beta surfaces drift), the platform fires a session autonomously at each scheduled time, keeps a per-firing run record, and lets you pause, unpause, or archive the schedule. Combined with [Agent Once, Session Every Run](/concepts/agent-once-session-every-run), the shape is clean: a stable versioned agent, a clock, and a trail of runs. ## Why it is a bigger deal than it sounds Cron is trivial. Cron for agents is not, because each firing needs everything a session needs: a sandbox, credentials, memory, output collection, and something watching for the run that hangs or fails. Teams that self-host end up building a scheduler, a job queue, and an observability layer around what began as one recurring task. The per-firing run record is the quietly important part: recurring unattended work is exactly where [silent failure](/concepts/silent-failure) hides, and a record per firing is the difference between "the report did not arrive" and knowing which run broke and what it did. Anthropic's own [launch-your-agent](https://github.com/anthropics/launch-your-agent) reference skill treats scheduling as the final step of shipping: it scopes a founder's agent, grades it, and "if it should run on a clock, puts it on a scheduled deployment." Running without you is the definition of done. ## The pattern in the field The published stories show recurring, clock-driven agent work becoming a product shape of its own: - Eve Legal runs nightly agents across a law firm's entire caseload, so, in its CEO's words, "When you wake up, you're approving legal work that has been done overnight." ([source](https://claude.com/customers/eve-legal)) - Tasklet's whole product is this shape: users describe recurring business automations and the platform reports "450,000 agent actions executed per day across all customer agents." ([source](https://claude.com/customers/tasklet)) - Brainlabs runs Managed Agents fired from a Notion board as work changes state, the event-driven sibling of the cron trigger. ([source](https://claude.com/customers/brainlabs)) ## What to pair it with A scheduled agent is unattended by definition, so the surrounding disciplines stop being optional: [Memory Stores](/concepts/memory-stores) so runs build on each other, [Outcomes](/concepts/outcomes) so each run has a definition of done, and [the Human Gate](/concepts/the-human-gate) on any action that commits irreversibly. The overnight-work pattern in the Eve Legal story is the model: the agent works the night shift, the human approves in the morning. ## Further Reading - [Agent Once, Session Every Run](/concepts/agent-once-session-every-run): the object model schedules fire against. - [Long-Running Work](/concepts/long-running-work): the sibling workload class. - [Silent Failure](/concepts/silent-failure): the risk profile of unattended recurrence. - [The Human Gate](/concepts/the-human-gate): morning approval of the night shift's work. --- # Self-Correction Mid-Task URL: https://buildonanthropic.com/concepts/self-correction # Self-Correction Mid-Task *An agent that checks its work against reality before handing it over, and fixes what it finds.* ![Hand-drawn line illustration in warm terracotta and cream: a loopy line-drawn hand holding a cream stylus redraws a long wobbly line at the point it veered off course, the corrected path continuing forward while the abandoned crooked segment trails away, a one-line profile face observing.](/img/pages/concepts-self-correction.jpg) --- ## What it is Self-correction is the loop inside the loop: generate, verify against something real (a test run, a rendered screen, a compiler), and revise before declaring done. The demand-side phrasing comes from Coasty in YC's Summer 2026 batch, which describes an agent that "catches when it goes off track and fixes itself mid-task." The key word is mid-task: the check happens before the handoff, not after the damage. ## The pattern in production The published stories converge on one principle: the verification signal comes from reality, not from the model re-reading its own text. - CircleCI's Chunk agent generates a fix, "validates it in a real CI pipeline sandbox, retries on failure, and delivers a green pull request." The check is an actual pipeline run, and the trigger has inverted too: "More than 4 in 5 agent tasks now triggered automatically at the point of failure." ([source](https://claude.com/customers/circleci)) - Replit's Agent 4 tests its own apps by looking at them, "self-testing via vision" against the rendered interface, an approach its story reports "runs 3x faster and 10x more cost-effectively than computer-use approaches" on the same workflow. ([source](https://claude.com/customers/replit)) - Classmethod made the check a process gate: Claude Code "performs mandatory AI self-reviews before pull requests," which its story pairs with an "Up to 80% decrease in time spent reviewing code" for the humans downstream. ([source](https://claude.com/customers/classmethod)) - Bolt's design-system agent runs a self-review loop across its roughly 53-minute autonomous runs, resolving conflicts across thousands of components before a human sees the output. ([source](https://claude.com/customers/bolt)) ## The boundary with grading Self-correction has a known failure mode: the model grading its own work inside the same context is biased toward believing it succeeded, which is why Anthropic's guidance separates generation from judgment (the evaluator-optimizer pattern in [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents), December 2024, and the fresh-context evaluator in its long-running-agents workshop material). The resolution is a division of labor. Mid-task self-correction runs against objective signals: tests pass or they do not, the page renders or it does not. Final acceptance belongs to a separate judge, whether that is [Outcomes](/concepts/outcomes) with its own grader context or a human at [the gate](/concepts/the-human-gate). An agent checking its work against reality is engineering; an agent taking its own word for it is [silent failure](/concepts/silent-failure) with extra steps. ## Why it matters for long runs The longer the run, the further off course an uncorrected error compounds. The corpus's hours-long autonomous runs (see [Long-Running Work](/concepts/long-running-work)) are viable precisely because verification is woven through them; Rakuten's seven-hour run and Replit's six-hour sessions are continuous loops of doing and checking, not seven hours of unchecked output. Self-correction is what converts model capability into run length. ## Further Reading - [Outcomes](/concepts/outcomes): final acceptance by a separate grader. - [Silent Failure](/concepts/silent-failure): what this pattern exists to prevent. - [Long-Running Work](/concepts/long-running-work): the workload it makes possible. --- # Silent Failure URL: https://buildonanthropic.com/concepts/silent-failure # Silent Failure *The failure mode that kills trust in unattended agents: the agent thinks it finished, and it did not.* ![Hand-drawn ink-line illustration on terracotta: a cream vessel quietly tipped over with a thin ink line spilling out, a relaxed loopy hand resting nearby, a one-line profile face unaware beside it.](/img/pages/concepts-silent-failure.jpg) --- ## The problem it names A loud failure pages someone. A silent failure ships: the run ends green, the report goes out, and the mistake is discovered downstream by whoever it hurt. For agents doing real unattended work this is the trust problem, and the founders building those agents describe it in exactly those terms. Coasty, in YC's Summer 2026 batch, puts it in its launch material: "The problem isn't getting an agent to click. It's trusting it when a silent mistake gets someone fired." And sharper still: "An agent that thinks it finished but didn't is worse than one that fails loudly." ## What it looks like in the published record - Twilio's Feature Factory story contains the corpus's cleanest specimen. Mid-build, its author discovered "The MCP server... had been silently failing since the day it was configured." The cost: "thirty-five sessions of building and testing, all without the tool infrastructure the whole system was designed around." Everything appeared to work; the system had been routing around a dead dependency the entire time. ([source](https://claude.com/customers/twilio-qa)) - Vambe's multi-agent sales platform lived the statistical version: "Previously, handoffs between agents failed 30% of the time. With Claude, reliability improved to over 95%, and the average number of agents per platform increased from 1.0 to 2.3." A 30 percent silent handoff loss is not an outage; it is a product quietly underperforming. ([source](https://claude.com/customers/vambe)) - Freedom Forever's permit agent surfaced the human baseline's own silent failures on arrival: "The agent found emails that had been sitting unprocessed for an average of three weeks." Silent failure predates agents; unattended agents just concentrate the risk. ([source](https://claude.com/customers/freedom-forever)) ## The engineering answers Anthropic's long-running-agents workshop material ([cwc-long-running-agents](https://github.com/anthropics/cwc-long-running-agents)) encodes the two structural defenses. First, a **default-FAIL contract**: a run counts as failed unless it affirmatively proves success, inverting the optimistic default that lets half-done work pass. Second, a **fresh-context evaluator**: completion is judged by a checker that did not do the work, because the worker's own context is contaminated by its belief that it finished. On Managed Agents, [Outcomes](/concepts/outcomes) is that second defense as a platform primitive, and one design goal of tool-error handling in agent harnesses cuts the other way: the loop retrying past errors is what keeps a run alive, and also what can paper over a dead tool, as the Twilio case shows. Instrumentation on tool success rates is the countermeasure. ## Why it earns a concept page Because it reorders the engineering priorities. Teams default to chasing capability (can the agent do it?) when the adoption blocker is verification (do we know it did?). The corpus's most trusted unattended deployments lead with verification machinery: rubric graders, [audit trails](/concepts/the-audit-trail), [human gates](/concepts/the-human-gate) at the commit points. Trust is not a model property; it is an architecture. ## Further Reading - [Outcomes](/concepts/outcomes): the platform's separate-grader answer. - [Self-Correction Mid-Task](/concepts/self-correction): catching drift before the run ends. - [The Audit Trail](/concepts/the-audit-trail): making verification cheap for the human. - [The Human Gate](/concepts/the-human-gate): the backstop when verification must be human. --- # The Audit Trail URL: https://buildonanthropic.com/concepts/the-audit-trail # The Audit Trail *Citation to source as a product feature: every answer carries a pointer to the exact place it came from, so the human verifies instead of hunts.* ![Hand-drawn line illustration in warm terracotta and cream: a loopy line-drawn hand holds up a single cream card joined by one continuous wobbly line back to the small stack of cream cards it came from, a one-line profile face watching.](/img/pages/concepts-the-audit-trail.jpg) --- ## What it is An audit trail, in the agent-product sense, is the discipline of making every output traceable: each extracted fact, each drafted claim, each triage decision points back to the source document, sentence, or record it came from. The founders building for regulated and high-stakes buyers treat it as table stakes; Alloovium, in YC's Summer 2026 batch, describes a system where "every answer cites its exact source sentence," paired with the policy that "nothing AI-drafted moves without a human sign-off." ## Why it changes the economics of review Citation converts review from search to spot-check. Without it, the human verifying an AI output must re-find the evidence, which can cost nearly as much as doing the work; with it, verification is a click. The corpus quantifies the difference: - Carta Healthcare's abstraction product answers clinical registry questions "with justifications and direct citations, so human abstractors validate rather than hunt," per its story, which also reports abstractors previously spending "on average, one hour per case manually combing through clinical notes and documents" and quality holding at "98-99% Inter-rater Reliability (IRR) scores." ([source](https://claude.com/customers/carta-healthcare)) - Quillit measured citation quality directly and published the delta: "Before using Claude 3.5 Sonnet, our accuracy rating for citations was around 60-70%. After that, it was between 89-98%." Its stated stake: precision is "crucial for our clients' confidence in the insights they receive." ([source](https://claude.com/customers/quillit)) - Elation Health ships cited summaries of a patient's longitudinal chart inside the EHR, and reports "Reduced median time to first insight by 61% for chart review." ([source](https://claude.com/customers/elation-health)) ## The other trail: what the agent did Product citations answer "where did this claim come from." Operating an agent adds a second question: "what did the run actually do?" The platform's primitives leave that trail as a by-product when used properly: per-firing run records on [Scheduled Deployments](/concepts/scheduled-deployments), versioned writes on [Memory Stores](/concepts/memory-stores), immutable config versions under [Agent Once, Session Every Run](/concepts/agent-once-session-every-run), and persisted approval decisions at [the Human Gate](/concepts/the-human-gate) (details in the [platform docs](https://platform.claude.com/docs); verify against platform.claude.com/docs, beta surfaces drift). Together they answer the auditor's question: which config, which run, which sources, who approved. ## Why it earns a concept page Because it is the feature that makes [the Human Gate](/concepts/the-human-gate) fast enough to keep. An approval step whose reviewer must hunt for evidence gets bypassed under deadline pressure; an approval step with citations attached stays cheap enough to survive contact with a busy Monday. Traceability is not compliance overhead on the product; it is what lets the product be trusted at speed. The same standard this wiki applies to published metrics (see [Strength of Evidence](/case-studies/strength-of-evidence)) applies inside an agent product: a claim you can check beats a claim you must take on faith. ## Further Reading - [The Human Gate](/concepts/the-human-gate): the approval step citations accelerate. - [Silent Failure](/concepts/silent-failure): the trust problem traceability addresses. - [Strength of Evidence](/case-studies/strength-of-evidence): the same principle applied to published numbers. --- # The Batch API URL: https://buildonanthropic.com/concepts/the-batch-api # The Batch API *50 percent off input and output tokens for anything that can wait up to 24 hours.* ![Hand-drawn line illustration in warm terracotta and cream: two loopy line-drawn hands push a wide cream tray packed with many small cream squares while a third hand pinches a single tiny square, a one-line profile face observing between them.](/img/pages/concepts-the-batch-api.jpg) --- ## What it is The Message Batches API takes a file of requests instead of one request at a time. Per the [platform docs](https://platform.claude.com/docs) (verify against platform.claude.com/docs; limits and pricing drift): batches run up to 100,000 requests, results arrive within 24 hours and usually much sooner, and both input and output tokens are billed at 50 percent of the standard price. The discount stacks with [prompt caching](/concepts/prompt-caching), so a cached, batched input token carries both reductions. ## The sorting question The Batch API turns one product question into a pricing lever: **does this call need an answer while someone is waiting?** Interactive chat does. Almost everything else in a real product does not: nightly enrichment, classification backfills, embedding-adjacent preprocessing, eval suite runs, report generation, content pipelines, retroactive re-processing after a prompt improvement. Any of those running on the synchronous API is paying double for latency nobody is experiencing. The corpus is full of workloads with exactly this shape, whether or not each story names the API surface it bills through: Scribd generates metadata across a corpus of documents at a scale it reports as "70% of content improved with AI-generated metadata" ([source](https://claude.com/customers/scribd)), and Local Falcon reports "1 million customer reviews analyzed simultaneously" ([source](https://claude.com/customers/local-falcon)). Bulk, non-interactive, tolerant of hours: the batch profile. ## Why it pairs with evals The least obvious high-value batch workload is your own eval suite. Anthropic's [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) (January 2026) pushes teams toward evaluating continuously, and eval runs are the perfect batch customer: large, repetitive, cache-friendly, and never blocking a user. Cutting the eval bill in half changes how often a team can afford to run them, which changes how fast the product improves. Eve Legal, for instance, reports "around 420 test suites run against every model the company uses." ([source](https://claude.com/customers/eve-legal)) ## The honest boundary Batch is a discount on patience, not a workaround for capacity planning. Results within 24 hours is the contract, so anything with a same-hour SLA stays synchronous, and a pipeline that feeds a human's morning review needs to be submitted with the window in mind, the way [scheduled, overnight agent work](/concepts/scheduled-deployments) already is. ## Further Reading - [Prompt Caching](/concepts/prompt-caching): the discount this one stacks with. - [The Effort Dial and Model Splitting](/concepts/the-effort-dial): the third cost lever. - [Consumption Pricing](/concepts/consumption-pricing): the pricing model all the levers live inside. - [Scheduled Deployments](/concepts/scheduled-deployments): the natural producer of batchable work. --- # The Company Brain URL: https://buildonanthropic.com/concepts/the-company-brain # The Company Brain *A persistent store of how a specific company works, which agents read on every run and write back to as they learn.* ![Hand-drawn ink-line illustration on terracotta: a large open head profile drawn in one continuous line, with a loopy hand placing small cream blocks inside it.](/img/pages/concepts-the-company-brain.jpg) --- ## What it is A company brain is organizational knowledge made machine-usable: the procedures, standards, vocabulary, and accumulated judgment of one specific company, held somewhere agents can read it at the start of every run and improve it at the end. It is what turns a generic model into something that works here, at this firm, the way this firm works. The term is not this wiki's coinage. YC's Summer 2026 Requests for Startups names "Company Brain" as a category, framed as "a living map of how a company works," and three companies in that batch (Rational, Kebra, and Levocred) use the exact phrase in their own launch descriptions. Rational, for one, describes agents that "onboard themselves to the firm, build a company brain." The demand side of this pattern is explicit. ## The supply side, in the published corpus The customer stories show what the pattern looks like shipped, usually as skills plus persistent memory: - Brainlabs turned its agency's know-how into "A library of roughly 400 skills authored by employees in four weeks," including an SEO audit skill its CTO describes as "one skill, refined across thousands of iterations from our SEO team, that captures the full universe of what to look for and what good looks like." ([source](https://claude.com/customers/brainlabs)) - YMCA South Australia built "20+ custom skills" that its story says were "built to encode organizational knowledge, brand standards, and operational procedures" across 65 locations. ([source](https://claude.com/customers/ymca-south-australia)) - Notion closes the write-back loop with a "Self-improving skills database" that is "maintained by Claude after every completed task," so each run's lessons are available to the next. ([source](https://claude.com/customers/notion-qa)) - Rakuten's specialist Managed Agents run with "persistent compute and memory" beside employees across engineering, product, sales, marketing, and finance. ([source](https://claude.com/customers/rakuten-qa)) ## The mechanics underneath On the Claude platform the pattern decomposes into pieces that already exist: Agent Skills carry the procedures (Anthropic's [Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) post, October 2025, is the canonical description), [Memory Stores](/concepts/memory-stores) carry the accumulating state, and [a stable versioned agent](/concepts/agent-once-session-every-run) carries the configuration. The distinctive engineering problem is curation: a brain that only ever appends becomes noise, so the write-back path needs the same review discipline as code. Notion's thumbs-up-and-down feedback loop on its skills database is one published answer. ## Why it compounds A company brain is the rare agent asset that appreciates. Models improve and [harnesses depreciate](/concepts/the-harness), but a well-curated store of how-we-work transfers to every better model that arrives, and every run that writes back makes the next run start further ahead. That is also the moat argument: two competitors can buy the same model, but not the same brain. ## Further Reading - [Memory Stores](/concepts/memory-stores): the persistence primitive underneath. - [Agent Once, Session Every Run](/concepts/agent-once-session-every-run): the config that carries the skills. - [Long-Running Work](/concepts/long-running-work): the runs the brain accumulates from. - [The Index](/case-studies/the-index): every published build, searchable by pattern. --- # The Deployment URL: https://buildonanthropic.com/concepts/the-deployment # The Deployment *Everything an agent runs on: sandboxing, session persistence, tenant isolation, secrets handling, and capacity. The infrastructure half of the agent question.* ![Hand-drawn ink-line illustration on terracotta: two loopy hands stack wide cream slabs into a low foundation with a small block resting on top, a one-line profile face observing.](/img/pages/concepts-the-deployment.jpg) --- ## What it is [The harness](/concepts/the-harness) is the loop; the deployment is the ground it stands on. Anthropic's [Scaling Managed Agents](https://www.anthropic.com/engineering/managed-agents) post (April 2026) draws the line between the two, and the distinction is the most useful sorting question on the platform, because two of [the four ways to build](/concepts/four-ways-to-build) hand you a harness while still leaving the entire deployment to you. ## The bill of materials The Agent SDK documentation at [code.claude.com](https://code.claude.com/docs/en/agent-sdk) makes the deployment burden concrete, because the SDK is harness-only and the docs enumerate what the builder still owns (verify against the current docs; specifics drift): - **Sandboxing.** The SDK spawns a subprocess that owns a shell and a working directory. Containing it is your job: Docker, gVisor, Firecracker, a cloud sandbox, or your own VMs. - **Session persistence.** Session transcripts live on local disk, so a container restart loses the session even if you saved the session ID. Durable sessions mean wiring your own store. - **Tenant isolation.** Per-tenant working directories, per-tenant configuration, egress rules. Nothing isolates customers from each other for you. - **Secrets.** Tool credentials visible to the agent's environment are credentials a prompt injection can go looking for. Keeping them out requires an egress proxy you build. See [Vaults](/concepts/vaults) for the managed answer. - **Capacity.** The docs' sizing guidance is on the order of a gigabyte of RAM and several gigabytes of disk per concurrent agent, with memory growing over a session's life. Autoscaling that is an infrastructure project. Anthropic's security posts describe what its own deployments do about the first item: [Claude Code sandboxing](https://www.anthropic.com/engineering/claude-code-sandboxing) (October 2025) covers filesystem and network isolation boundaries, and [How we contain Claude across products](https://www.anthropic.com/engineering/how-we-contain-claude) (2026) covers the broader containment model. ## Why it matters For a small team, this list is the real cost of an agent product. The model call is minutes of work; the deployment is months of it. That is the trade [the four ways to build](/concepts/four-ways-to-build) makes visible: Managed Agents is the only path where Anthropic supplies the deployment, which is exactly why its customer stories talk about calendar time. ## How it shows up in the field - Sentry shipped its Managed Agents integration "in weeks instead of months" with a single engineer, and credits the platform with removing "the ongoing operational overhead of maintaining bespoke agent infrastructure." ([source](https://claude.com/customers/sentry)) - Pendo spent about three months trying to build its agent infrastructure in-house, then reached production on Managed Agents in "About three days." ([source](https://claude.com/customers/pendo-qa)) - Greptile, a ten-engineer team, frames the same trade from the SDK side: "For a company reviewing over a billion lines of code each month, the ability to focus on domain expertise rather than infrastructure has become a strategic advantage." ([source](https://claude.com/customers/greptile)) ## Further Reading - [The Harness](/concepts/the-harness): the loop that runs on top of all this. - [The Four Ways to Build](/concepts/four-ways-to-build): which path leaves you holding which parts. - [Vaults](/concepts/vaults): the managed answer to the secrets problem. - [The Startup Cut](/case-studies/startups): small teams and what they chose to own. --- # The Effort Dial and Model Splitting URL: https://buildonanthropic.com/concepts/the-effort-dial # The Effort Dial and Model Splitting *Spend frontier-model tokens on judgment and small-model tokens on commodity steps, and turn the effort dial down where deep thinking is wasted.* ![Hand-drawn line illustration in warm terracotta and cream: a loopy line-drawn hand turns a large plain cream dial with no markings while three cream blocks of increasing size sit in a row beside it, a one-line profile face watching.](/img/pages/concepts-the-effort-dial.jpg) --- ## Two levers, one idea Not every step in an agent's work deserves the same intelligence. The platform exposes that as two levers: - **The effort dial.** The Messages API takes an effort setting (levels from low up through the maximum) that trades thoroughness for speed and tokens; lower effort means fewer, more consolidated tool calls and less preamble. Per the [platform docs](https://platform.claude.com/docs) (verify against platform.claude.com/docs; parameter names and levels drift), high effort suits agentic coding and hard reasoning, low effort suits subagents and simple steps. - **Model splitting.** Route each step to the cheapest model that clears its quality bar: a frontier model for judgment calls, a small fast model for extraction, classification, and formatting. Anthropic's post [The advisor strategy](https://claude.com/blog/the-advisor-strategy) (April 2026) formalizes one version: a more capable model advises and plans, a cheaper model executes. The same logic gates whole architectures. [Building multi-agent systems: when and how](https://claude.com/blog/building-multi-agent-systems-when-and-how-to-use-them) (January 2026) reports that "multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks," which makes multi-agent a spend you justify, not a default. ## The split in production - Elation Health runs its clinical chart summaries on Claude Haiku 4.5, the small fast model, because in-visit latency is the constraint, and reports "Reduced median time to first insight by 61% for chart review." ([source](https://claude.com/customers/elation-health)) - Eve Legal makes Opus its production default for legal judgment and "uses Claude Sonnet for lighter extraction work" across the 12.5 million documents it processes a month. ([source](https://claude.com/customers/eve-legal)) - Replit splits by difficulty inside one product: Sonnet handles sustained development work while Opus takes architectural decisions and complex multi-file refactoring. ([source](https://claude.com/customers/replit)) - Asana states the principle outright: "By using different Claude models, we're able to optimize for speed, reasoning power, or balance between the two." ([source](https://claude.com/customers/asana-qa)) ## How to find your split Work backward from an eval, not forward from a price sheet. Run each pipeline step against the smallest model at low effort, measure, and promote only the steps that fail. Teams that skip the measurement default everything to the frontier model, which is the expensive kind of caution. The eval harness that [Outcomes](/concepts/outcomes) and [Strength of Evidence](/case-studies/strength-of-evidence) both argue for is the same harness that makes downgrading safe. ## Further Reading - [Prompt Caching](/concepts/prompt-caching) and [The Batch API](/concepts/the-batch-api): the other two cost levers; all three stack. - [Consumption Pricing](/concepts/consumption-pricing): the economics that make these levers matter. - [Outcomes](/concepts/outcomes): the eval loop that tells you which steps can afford a smaller model. --- # The Harness URL: https://buildonanthropic.com/concepts/the-harness # The Harness *The software loop around the model: tool dispatch, context management, retries, and compaction. The half of an agent that is code rather than model.* ![Hand-drawn ink-line illustration on terracotta: a loopy hand traces a continuous looping circuit around a single cream block, a one-line profile face observing.](/img/pages/concepts-the-harness.jpg) --- ## What it is An agent is a model plus a loop. The model decides; the loop does everything else: it dispatches tool calls, feeds results back, manages the context window, retries failures, and compacts the conversation when it grows too long. Anthropic's engineering post [Scaling Managed Agents](https://www.anthropic.com/engineering/managed-agents) (April 2026) names that loop the harness, and separates it cleanly from [the deployment](/concepts/the-deployment), the infrastructure the loop runs on. The concrete jobs a harness does, drawn from Anthropic's own harness writing: run the tool-use loop, decide what stays in context, apply [prompt caching](/concepts/prompt-caching), trigger compaction as the context window fills, and keep a long task on the rails with progress files and checkpoints. The engineering post [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) (November 2025) adds the long-horizon toolkit: initializer agents, progress files, environment management. ## Why it matters Anthropic's central claim about harnesses is that they age badly. The blog post [Agent harness design](https://claude.com/blog/harnessing-claudes-intelligence) (April 2026) puts it directly: "Agent harnesses encode assumptions about what Claude can't do on its own, but those assumptions grow stale as Claude gets more capable." Scaffolding written to compensate for last year's model becomes drag on this year's. The Managed Agents post makes the same argument the reason for the product's design: harness assumptions go stale, so the platform is built around interfaces that stay stable while harnesses change underneath. The practical consequence: treat a harness as a depreciating asset. Code you wrote to babysit the model should be deleted as models improve, not defended. ## Who supplies it Every one of [the four ways to build](/concepts/four-ways-to-build) is an answer to two questions, and this is the first one: who supplies the harness. Write a manual loop over the Messages API and you own it. Use Tool Runner or the Agent SDK and the SDK owns it. Use Managed Agents and Anthropic runs it as a service, with compaction and caching included by default (per the [Managed Agents docs](https://platform.claude.com/docs); verify against platform.claude.com/docs, beta surfaces drift). ## How it shows up in the field - CircleCI built Chunk, its task-to-validated-pull-request agent, on the Agent SDK's harness rather than its own: "We could not have built Chunk without the Claude Agent SDK," per its [customer story](https://claude.com/customers/circleci). - Greptile runs its code-review agent on the Agent SDK and reports "~90% cache hit rates" from the harness's caching behavior. ([source](https://claude.com/customers/greptile)) - Sentry moved its fix-writing agent onto Managed Agents so a single engineer could ship the integration, saying the platform "eliminated the ongoing operational overhead of maintaining bespoke agent infrastructure." ([source](https://claude.com/customers/sentry)) ## Further Reading - [The Deployment](/concepts/the-deployment): the other half of the split, the infrastructure the loop runs on. - [The Four Ways to Build](/concepts/four-ways-to-build): who supplies the harness and deployment on each path. - [Long-Running Work](/concepts/long-running-work): the workload class that stresses harnesses hardest. - [Prompt Caching](/concepts/prompt-caching): the harness behavior with the largest cost consequence. --- # The Human Gate URL: https://buildonanthropic.com/concepts/the-human-gate # The Human Gate *An approval step the platform enforces before an agent's action commits. A prompt can ask for approval; a permission policy guarantees it.* ![Hand-drawn ink-line illustration on terracotta: a loopy hand raised calmly palm-forward inside a simple cream archway, a one-line profile face observing beside it.](/img/pages/concepts-the-human-gate.jpg) --- ## What it is A human gate is a checkpoint where an agent must stop and get sign-off before an action takes effect: before the filing submits, the email sends, the money moves. On Managed Agents this is a platform feature: per the [docs](https://platform.claude.com/docs) (verify against platform.claude.com/docs; beta surfaces drift), permission policies mark each tool `always_allow` or `always_ask`, and an `always_ask` tool triggers a confirmation round-trip that the harness enforces. The Agent SDK has the same machinery in permission modes and hooks, per [its docs](https://code.claude.com/docs/en/agent-sdk). ## Enforcement beats instruction The load-bearing distinction: telling the model "always ask before submitting" is an instruction, and instructions are probabilistic. The model follows them almost always, and almost always is not a compliance posture. A permission policy is enforcement; the tool call cannot execute without the approval, no matter what the model decides, and no matter what a prompt injection tells it to decide. That last clause is why the gate is also a security control and the standing complement to [Vaults](/concepts/vaults): the vault protects credentials, the gate governs actions. ## The demand is explicit The builders this platform courts are asking for exactly this. In YC's Summer 2026 batch launch material, Zomma states it plainly: "Every action stops for human approval before anything submits." Alloovium, in the same batch, ships the same posture: "nothing AI-drafted moves without a human sign-off." Anthropic's own advice in [Building AI agents for startups](https://claude.com/blog/building-ai-agents-for-startups) points the same direction: start "in areas where human oversight already exists and imperfect automation won't create major problems." ## The gate in production - Duvo's procurement agents operate SAP and supplier portals for multi-billion-euro retailers with human-in-the-loop approval, and persist each approval decision so the same question is not asked twice. ([source](https://claude.com/customers/duvo)) - Eve Legal's overnight agents end at a gate by design: "When you wake up, you're approving legal work that has been done overnight." ([source](https://claude.com/customers/eve-legal)) - Notion routes finished agent work to a named approver: "our platform routes them to the right person for approvals." ([source](https://claude.com/customers/notion-qa)) - Twilio's Feature Factory story shows the gate as a dial rather than a switch: "On day one, I had human approval gates after every pipeline phase. Six interruptions per feature. By day 45, the system ran headless: no human in the loop, real phone calls, real validation." ([source](https://claude.com/customers/twilio-qa)) The Twilio arc is the mature version of the pattern: gates everywhere at first, then removed one at a time as validation earns the removal, never assumed away up front. ## Further Reading - [Vaults](/concepts/vaults): the credential half of the security story. - [Silent Failure](/concepts/silent-failure): what unattended agents do without a gate. - [The Audit Trail](/concepts/the-audit-trail): what the approver needs in order to approve fast. - [Scheduled Deployments](/concepts/scheduled-deployments): the overnight-work pattern the gate completes. --- # The No-API Integration URL: https://buildonanthropic.com/concepts/the-no-api-integration # The No-API Integration *Operating the screens of systems that will never get an API, with computer use standing in for the integration nobody will ever ship.* ![Hand-drawn line illustration in warm terracotta and cream: a loopy line-drawn hand presses a fingertip directly onto the blank cream face of a plain rectangular panel while a cable ends nearby in a connector plugged into nothing, a one-line profile face watching.](/img/pages/concepts-the-no-api-integration.jpg) --- ## The problem it names A large share of real business workflows terminate in software with no API: utility permit portals, legacy ERPs, supplier extranets, government filing systems. The demand signal is unusually crisp: in a full pull of YC's Summer 2026 batch, 27 percent of companies describe integrating with API-less systems, the largest architecture requirement in the batch. Zomma's launch material names the category, "systems that will never get an API," and Whitespace describes the status quo it replaces: "6 reps typing 600 lines of orders from email into the ERP every day." The move is to stop waiting for the API and operate the screen: the agent sees the interface, clicks, types, and reads the result, the way the human it replaces did. ## Proven in the corpus - Freedom Forever automates solar permit submissions on third-party utility websites with the Agent SDK, reporting a "97.5% success rate on complex web automation benchmarks where alternative frameworks hit 20–60%," built against eight internal mock sites run five iterations each. It also reports "~19,000 permit emails processed" in the first month of production. ([source](https://claude.com/customers/freedom-forever)) - Duvo's agents "operate real enterprise interfaces" for multi-billion-euro retailers, per its story: SAP GUIs, supplier portals, spreadsheets, email, phone. One agent run might "log into a supplier portal, extract delivery status for 50 purchase orders, cross-reference against SAP." Reported result: "€2.8M+ in annualized savings" for one customer in three months. ([source](https://claude.com/customers/duvo)) - Tasklet reaches the long tail the same way: alongside 3,000+ prebuilt integrations, "Claude spins up a browser in a cloud VM" when no integration exists. ([source](https://claude.com/customers/tasklet)) ## What makes it work in practice Two disciplines separate the production stories from the demos. First, benchmarks before deployment: Freedom Forever built mock sites replicating real-world failure modes and measured against them, which is why its success number carries weight (see [Strength of Evidence](/case-studies/strength-of-evidence)). Second, containment around the click: an agent driving a real screen with real credentials is the maximum-blast-radius configuration, which is why the pattern pairs with [Vaults](/concepts/vaults) for the credentials and [the Human Gate](/concepts/the-human-gate) before anything submits. Zomma's own posture, "Every action stops for human approval before anything submits," is the pattern stated as policy. ## The honest boundary Screen operation is slower and less deterministic than an API call, and always will be. Where an API exists, use it; MCP is the platform's native shape for that, per the [platform docs](https://platform.claude.com/docs). The no-API integration is for the workflows where the alternative is not a cleaner integration but a human keying data into a portal, and its economics are judged against that human baseline. Freedom Forever's published cost of "$0.20–$0.35" per agent run is the kind of number that settles the comparison. ([source](https://claude.com/customers/freedom-forever)) ## Further Reading - [Vaults](/concepts/vaults): keeping credentials out of the sandbox that drives the screen. - [The Human Gate](/concepts/the-human-gate): approval before the submit button. - [Silent Failure](/concepts/silent-failure): the risk profile of unattended screen work. - [The Index](/case-studies/the-index): more computer-use builds across the corpus. --- # Vaults URL: https://buildonanthropic.com/concepts/vaults # Vaults *The Managed Agents credential store: secrets never enter the sandbox. They are substituted into outbound requests at the egress boundary, so the agent never sees them.* ![Hand-drawn line illustration in warm terracotta and cream: a loopy line-drawn hand reaches toward a cream key that sits sealed inside a wobbly enclosed box, never crossing the boundary, a one-line profile face watching from above.](/img/pages/concepts-vaults.jpg) --- ## What it is A Vault holds the credentials an agent's tools need: API keys, OAuth tokens, environment-variable secrets. Per the [Managed Agents docs](https://platform.claude.com/docs) (verify against platform.claude.com/docs; beta surfaces drift), the design rule is that secrets never enter the sandbox where the agent runs. When the agent makes an outbound request that needs a credential, the platform substitutes the secret into the request at egress. OAuth tokens for MCP connections refresh automatically, and secrets can be scoped to an allowlist of hosts. ## Why the egress design matters Prompt injection is the standing threat against any agent that reads untrusted content: a hostile document or web page instructs the model to do something its operator did not intend. The worst version is exfiltration, where the injected instruction is "send your API key to this address." Egress substitution removes the target rather than defending it. The model cannot leak a credential it never held; there is no key in the context window, no token in an environment variable inside the sandbox, nothing for an injected instruction to find. An attacker who fully controls the model's output still cannot read a secret that only exists on the other side of the egress boundary, and host allowlisting bounds where a substituted credential can even be sent. Contrast this with what the self-hosted paths require. On the Agent SDK, per its [docs](https://code.claude.com/docs/en/agent-sdk), credentials the agent's tools use are visible to the agent's environment unless the team builds its own egress proxy, which is real infrastructure. This is one of the sharpest single items on [the deployment](/concepts/the-deployment) bill of materials. ## The boundary, honestly stated A Vault protects the credential, not the action. An injected instruction can still tell the agent to misuse a capability it legitimately holds, such as sending a request it should not send to an allowed host. That class of attack is what [the Human Gate](/concepts/the-human-gate) exists for, and Anthropic's security writing, including [How we contain Claude across products](https://www.anthropic.com/engineering/how-we-contain-claude) (2026), treats containment as layered for exactly this reason. Vaults close the exfiltration channel; approval policies govern the actions. ## Who feels this problem Teams operating agents against real business systems with real credentials. Duvo's agents log into supplier portals and operate SAP for multi-billion-euro retailers ([source](https://claude.com/customers/duvo)), and Tasklet's customers run unattended automations where "Claude spins up a browser in a cloud VM" against live accounts ([source](https://claude.com/customers/tasklet)). At that point credential handling stops being a checklist item and becomes architecture. ## Further Reading - [The Deployment](/concepts/the-deployment): where secrets sit in the full infrastructure bill. - [The Human Gate](/concepts/the-human-gate): the enforcement layer for actions rather than credentials. - [The No-API Integration](/concepts/the-no-api-integration): the pattern that puts agent credentials on third-party screens. --- # Glossary URL: https://buildonanthropic.com/reference/glossary # Glossary *One-line-per-term summary across the whole wiki. The lexicon, flattened.* --- Every coined term on one line, linked to its canonical page in [Concepts](/concepts). - [Agent Once, Session Every Run](/concepts/agent-once-session-every-run): create the versioned agent config once, create a session per run; agent creation in the request path is the anti-pattern. - [The Audit Trail](/concepts/the-audit-trail): citation to source as a product feature, so the human verifies instead of hunts. - [The Batch API](/concepts/the-batch-api): 50 percent off input and output tokens for anything that can wait up to 24 hours; stacks with caching. - [The Company Brain](/concepts/the-company-brain): a persistent store of how a specific company works, which agents read on every run and write back to as they learn. - [Consumption Pricing](/concepts/consumption-pricing): the platform bills by tokens consumed, not seats sold, so the bill scales with how much the product actually runs. - [The Deployment](/concepts/the-deployment): everything an agent runs on: sandboxing, session persistence, tenant isolation, secrets handling, and capacity. - [The Effort Dial and Model Splitting](/concepts/the-effort-dial): frontier-model tokens for judgment, small-model tokens for commodity steps, effort turned down where thinking is wasted. - [The Four Ways to Build](/concepts/four-ways-to-build): manual loop, Tool Runner, Agent SDK, Managed Agents, separated by who supplies the harness and who supplies the deployment. - [The Harness](/concepts/the-harness): the software loop around the model: tool dispatch, context management, retries, and compaction. - [The Human Gate](/concepts/the-human-gate): an approval step the platform enforces before an action commits; a prompt asks, a permission policy guarantees. - [Long-Running Work](/concepts/long-running-work): agent runs measured in hours rather than seconds: the workload class that makes every architecture decision matter. - [Memory Stores](/concepts/memory-stores): workspace-scoped persistent memory mounted into every session, versioned and auditable. - [The No-API Integration](/concepts/the-no-api-integration): operating the screens of systems that will never get an API. - [Outcomes](/concepts/outcomes): define done as a rubric, and the platform runs an iterate, grade, revise loop with a separate grader until the work passes. - [Prompt Caching](/concepts/prompt-caching): cached input reads cost a fraction of base price, but one changed byte in the prefix invalidates everything after it. - [Scheduled Deployments](/concepts/scheduled-deployments): cron for agents: a schedule fires a fresh session on the clock, and every firing leaves a run record. - [Self-Correction Mid-Task](/concepts/self-correction): an agent that checks its work against reality before handing it over, and fixes what it finds. - [Silent Failure](/concepts/silent-failure): the failure mode that kills trust in unattended agents: the agent thinks it finished, and it did not. - [Vaults](/concepts/vaults): secrets never enter the sandbox; they are substituted into outbound requests at egress. --- # Graphic Style URL: https://buildonanthropic.com/reference/graphic-style # Graphic Style *Visual conventions for illustrations and diagrams on this wiki. The default is comic-strip article summaries in the Supersuit Up family style.* --- This wiki is part of the [Supersuit Up wiki family](https://supersuit.wiki) and shares its visual identity. Articles open with a comic-strip hero in the same neo-comic action-zine style used across the family. ## Article-hero comic strips (default) The default illustration for any article on this wiki is a **comic-strip summary** in neo-comic action-zine style with high-tech power-armor visual DNA — the same family used across [supersuit.wiki](https://supersuit.wiki) and the rest of the Supersuit Up wiki family. Cream paper, bold inked outlines, crimson + cobalt + gold + arc-reactor cyan, first-person helmet-visor POV as the signature panel, panel count matched to the article's beats. **Canonical spec:** [supersuit.wiki/reference/graphic-style](https://supersuit.wiki/reference/graphic-style). The full prompt template, palette, panel-count rules, and IP-name discipline live there. **Generation workflow:** use the global `comic-strip` skill at `~/.agents/skills/comic-strip/SKILL.md`. It bundles two canonical reference images that **must** be passed via `--input-image` on every generation to lock the style. Once this wiki has 2+ comics of its own in the family style, use that wiki's own comics as references for in-wiki style continuity. ## Image embedding Save generated comics to `static/img/comics/.png` where `` matches the article's URL slug. Embed at the top of the article, immediately after the subtitle italic block, before the first `## section`: ```markdown *Article subtitle goes here.* ![One-line alt text describing the strip's beats.](/img/comics/.png) --- ## First section ``` --- ## Further Reading - [supersuit.wiki: graphic style](https://supersuit.wiki/reference/graphic-style): the canonical spec for the comic-strip style across the family - [Voice rules](/reference/voice-rules): the prose equivalent of this page --- # Reference Index URL: https://buildonanthropic.com/reference # Reference Index *The map to Anthropic's published material for builders. Every entry links to the original, which is always the authority.* ![Hand-drawn line illustration in warm terracotta and cream: a hand tips one cream book out of a neat shelf row of book spines, considered by an abstract one-line face.](/img/pages/reference.jpg) --- ## How this index works Anthropic's builder material lives on five separate surfaces that do not cross-reference each other in any organized way. This page lists what exists and where, so you can go to the primary source. Entries are summarized in one line and linked. Nothing is reproduced here. Where a date is known it is given, because in this field a post from eighteen months ago and a post from last month can describe genuinely different APIs. ## The five surfaces | Surface | Scale | What it actually holds | |---|---|---| | [anthropic.com/engineering](https://www.anthropic.com/engineering) | ~25 posts | The deep technical canon. Small, dense, high signal. | | [claude.com/blog](https://claude.com/blog) | ~193 posts | The primary developer stream. Largest and least discovered. Older `anthropic.com/news` developer URLs now redirect here. | | [platform.claude.com/docs](https://platform.claude.com/docs) | ~541 pages | Messages API, Managed Agents, prompting and evaluation guidance. | | [code.claude.com/docs](https://code.claude.com/docs) | ~171 pages | A separate domain. The entire Agent SDK reference lives here. | | Code with Claude sessions and talks | ~140 sessions, ~79 videos | Conference material, much of it with recordings. | | [Anthropic Academy](https://anthropic.skilljar.com/) | ~21 courses | Structured courses. | ## Start here: the load-bearing posts If you read nothing else, read these. Each is linked to the original. - [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) (Dec 2024). The founding taxonomy: the augmented LLM, then workflow patterns, then agents. The source of the recurring advice to find the simplest thing that works before adding complexity. - [Effective context engineering for AI agents](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) (Sep 2025). Treats context as a finite resource with diminishing returns. - [Writing effective tools for agents](https://www.anthropic.com/engineering/writing-tools-for-agents) (Sep 2025). Tool design as the agent-computer interface. - [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) (Jan 2026). How to evaluate when the output is a trajectory rather than a string. - [Scaling Managed Agents](https://www.anthropic.com/engineering/managed-agents) (Apr 2026). Anthropic's own framing of why harnesses go stale and what a managed platform stabilizes. - [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) (Nov 2025). Initializer agents, progress files, environment management. ## For founders specifically - [The founder's playbook: Building an AI-native startup](https://claude.com/blog/the-founders-playbook) (May 2026). Stage-by-stage, with named failure modes. - [Building AI agents for startups](https://claude.com/blog/building-ai-agents-for-startups) (Nov 2025). Where to start when oversight already exists and imperfect automation is survivable. ## Cost and unit economics The highest-ROI reading in the whole corpus for anyone running a real bill. - [Lessons from building Claude Code: prompt caching is everything](https://claude.com/blog/lessons-from-building-claude-code-prompt-caching-is-everything) (Apr 2026). Caching treated as architecture rather than optimization. - [Prompt caching with Claude](https://claude.com/blog/prompt-caching). The durable economics of cache writes versus cache reads. - [The advisor strategy](https://claude.com/blog/the-advisor-strategy) (Apr 2026). Pairing a more capable advisor model with a cheaper executor. ## Agent architecture and coordination - [Common workflow patterns for AI agents](https://claude.com/blog/common-workflow-patterns-for-ai-agents-and-when-to-use-them) (Mar 2026). Sequential, parallel, evaluator-optimizer, and when each is warranted. - [Building multi-agent systems: when and how](https://claude.com/blog/building-multi-agent-systems-when-and-how-to-use-them) (Jan 2026). Includes the token-cost multiple of multi-agent versus single-agent, and named failure modes. - [Multi-agent coordination patterns](https://claude.com/blog/multi-agent-coordination-patterns) (Apr 2026). Five approaches with where each struggles. - [How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system) (Jun 2025). - [Agent harness design](https://claude.com/blog/harnessing-claudes-intelligence) (Apr 2026). ## Tools, MCP, and skills - [Introducing advanced tool use](https://www.anthropic.com/engineering/advanced-tool-use) (Nov 2025). Tool search and programmatic tool calling. - [Code execution with MCP](https://www.anthropic.com/engineering/code-execution-with-mcp) (Nov 2025). Why writing code to call tools scales better than direct calls. - [Agent Skills](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) (Oct 2025). - [Contextual Retrieval](https://www.anthropic.com/engineering/contextual-retrieval) (Sep 2024). ## Security and containment - [How we contain Claude across products](https://www.anthropic.com/engineering/how-we-contain-claude) (2026). - [Claude Code sandboxing](https://www.anthropic.com/engineering/claude-code-sandboxing) (Oct 2025). Filesystem and network isolation boundaries. - [Zero Trust for AI agents](https://claude.com/blog/zero-trust-for-ai-agents) (May 2026) and [the agent identity access model](https://claude.com/blog/agent-identity-access-model) (Jun 2026). ## Runnable code - [anthropics/claude-cookbooks](https://github.com/anthropics/claude-cookbooks). Classification, retrieval, tool use, evaluations, caching. - [anthropics/claude-quickstarts](https://github.com/anthropics/claude-quickstarts). Complete starter projects. - [anthropics/launch-your-agent](https://github.com/anthropics/launch-your-agent). A Claude Code skill that interviews a founder, scopes a first version, launches it on Managed Agents, grades it, and schedules it. The closest thing Anthropic ships to a guided path from idea to running agent. - [anthropics/cwc-long-running-agents](https://github.com/anthropics/cwc-long-running-agents). Workshop material on long-horizon reliability. - [anthropics/skills](https://github.com/anthropics/skills) and [anthropics/courses](https://github.com/anthropics/courses). ## Where the gaps are Worth knowing what is not well covered, so you stop searching for it: unit economics of an agent product beyond token pricing, multi-tenant SaaS architecture, reliability engineering against the API (retries, idempotency, circuit breakers, SLOs), model portability, evals as an operational asset at scale, regulated-industry engineering patterns, and latency engineering. Anthropic also rarely publishes the case for not building an agent at all. ## Further Reading - [Case Studies](/case-studies) who has already shipped on this - [The Index](/case-studies/the-index) all 256 published customer stories in one table - [The startup cut](/case-studies/startups) companies that were small when they started - [Strength of evidence](/case-studies/strength-of-evidence) which published numbers hold up --- # Tools URL: https://buildonanthropic.com/reference/tools # Tools *One page per tool. What it is, what it costs, when to use it, and when not.* --- Every tool page in this wiki follows the same shape: - **What it is** — one paragraph, plain. - **What it costs** — pricing tiers, the honest version. - **When to use it (and when not)** — paired bullets. The "not" half is load-bearing. - **Setup pointer** — link to official docs. We do not duplicate installation steps. See `templates/tool.mdx` at the repo root for the copy-and-rename scaffold. --- # Voice Rules URL: https://buildonanthropic.com/reference/voice-rules # Voice Rules *Writing constraints every page is held to. Edit per wiki to match the operator's house style.* --- Voice rules keep the wiki coherent across contributors. Replace the defaults below with your wiki's specific constraints. ## Page anatomy (non-negotiable) - Frontmatter with `title`, `slug`, `description`, `sidebar_position` (where applicable). - H1 matches the title. - Italic one-line definition directly under the H1. - `---` divider before the body. - 3-5 named H2 sections in the body. - "Further Reading" section at the bottom with cross-links. ## Voice (edit per wiki) - Direct, economical. Every word earns its place. - Concrete over abstract. Specifics over categories. - No em dashes. Use periods, commas, or colons instead. - No motivational poster language. No clichés. - Define every coined term in [Concepts](/concepts). Cross-link, do not redefine. ## Cross-linking - Use Docusaurus-style absolute paths: `/concepts/term-name`, not relative. - Every page should link to at least two other pages in the Further Reading section. - Concepts link to each other. The lexicon is a graph. --- # Start Here URL: https://buildonanthropic.com/start-here import ChangelogWidget from '@site/src/components/ChangelogWidget'; # Build on Anthropic *Anthropic's guidance for builders is excellent and scattered across five separate surfaces with no unified index. This is the index.* ![Hand-drawn line illustration in warm terracotta and cream: two hands stack cream blocks into a tower on a wide shared foundation slab, beside a receding row of smaller block towers, while an abstract face drawn in a single line looks on.](/img/hero-building-on-the-foundation.jpg) --- ## The problem this solves If you are a founder deciding whether and how to build your product on Claude, the material you need already exists. Finding it is the hard part, because it is spread across five places that do not link to each other in any organized way: | Surface | Roughly how much | What most people assume | |---|---|---| | `anthropic.com/engineering` | 25 posts | That this is the canon. It is the smallest slice. | | `claude.com/blog` | 193 posts | Largely undiscovered. This is the actual primary developer stream. | | `platform.claude.com/docs` | 541 pages | Reference, not narrative. | | `code.claude.com/docs` | 171 pages | A separate domain. The entire Agent SDK lives here. | | Conference sessions and talks | ~140 sessions, 79 videos | Rarely surfaced at all. | | Anthropic Academy | 21 courses | Rarely surfaced at all. | A reader who finds only the engineering blog misses roughly 85 percent of it. The same is true of the proof. `claude.com/customers` publishes 256 customer stories, and its own index page lazy-loads a fraction of them at a time. Roughly 95 more customer stories live entirely outside that index, scattered through news posts, blog posts, and conference sessions. ## How to use this wiki Start with the question you actually have. - **"Has anyone like me built this already?"** Go to [Case Studies](/case-studies). 256 stories, cut by company stage, industry, use-case pattern, and how hard the published evidence is. [The Index](/case-studies/the-index) lists every story in one searchable table. If you are early, start with [the startup cut](/case-studies/startups). - **"What has Anthropic actually published?"** Go to [the Reference Index](/reference), the dated master list of every source, organized by surface. - **"What should I build, concretely?"** Go to [The Tools](/tools): paste one file into Claude and get a personalized architecture read on your own product, ending in a build brief. Operators can mint custom editions for their own audience. Every page here is a routing layer first. It summarizes, then points you at the original. The originals are always linked and always the authority. ## For agents and LLMs This wiki ships machine-readable versions of itself, and they are meant to be used: - [`/llms.txt`](https://buildonanthropic-wiki.vercel.app/llms.txt) is the index: every page, one line each, with URLs. - [`/llms-full.txt`](https://buildonanthropic-wiki.vercel.app/llms-full.txt) is the entire wiki as one plain-text file, including [The Index](/case-studies/the-index)'s full 256-row case-study table. Point a coding agent, a research assistant, or a build brief at either file and it can read everything here without scraping. Both regenerate on every deploy, so they are always current. ## How claims are handled This wiki is source-grounded. That means: - Every claim names its source in the sentence. - Every number is the figure the company or Anthropic published, quoted briefly and linked to the page it came from. - Where a published source contradicts itself, that is flagged rather than smoothed over. - Nothing here is reproduced wholesale. This is a map to the material, not a copy of it. If a number matters to a decision you are making, click through and read the original. ## Independent, not affiliated **This wiki is an independent reference maintained by Gary Sheng. It is not affiliated with, endorsed by, or an official publication of Anthropic.** Anthropic's own documentation is the authority on every technical detail. Where this wiki and Anthropic's documentation disagree, Anthropic is right and this is out of date. ## Changelog For the full month-by-month log of every change, see the [Wiki Changelog](/changelog). --- # The Tools URL: https://buildonanthropic.com/tools # The Tools *Two plain markdown files you paste into Claude. One generates a personalized architecture deck for a founder. The other generates custom editions of the generator itself.* ![Hand-drawn line illustration in warm terracotta and cream: a loopy hand presses an ink stamp onto a cream sheet while identical stamped sheets recede to the right, watched by an abstract face drawn in a single line.](/img/pages/tools.jpg) --- ## 1. Build on Claude: the personalized architecture read **File: [`/tools/build-on-claude.md`](https://www.buildonanthropic.com/tools/build-on-claude.md)** For any founder or builder, in an accelerator or not. Copy the file, paste it into Claude (best: Claude Code inside your repo, so it reads your context first), and it interviews you for the gaps, then generates a clickable deck personalized to your product: where you sit among the four ways to build, working code for your case, honest cost arithmetic, the trade-offs including when NOT to build on Claude, and a copyable build brief that starts your first real build session. It is honest by construction: if the right answer for you is one API call and no agent, the deck says so. If staying provider-portable is right for your situation, the deck says that too. ## 2. Customize the generator: mint your own edition **File: [`/tools/customize-the-generator.md`](https://www.buildonanthropic.com/tools/customize-the-generator.md)** For evangelists, accelerator partners, community leads, and agencies. Paste it into your harness and it interviews you about your audience (a batch, a vertical, a community), pulls the most relevant proof rows from [The Index](/case-studies/the-index), and emits a customized edition of the generator carrying your name, your audience tuning, and your call to action. The honesty bar carries through every edition and cannot be customized away. A generator that flatters its audience burns the operator's name with it. ## How the pieces fit The wiki is the evidence layer: [The Index](/case-studies/the-index) holds all 256 published builds, and [`/llms-full.txt`](https://www.buildonanthropic.com/llms-full.txt) carries the whole wiki as one machine-readable file. The tools are the action layer: they read that evidence at generation time, so every deck's proof slide quotes real published numbers, never inventions. Paste chain, end to end: an operator customizes the generator, a founder runs their edition and gets a deck, the deck hands them a build brief, the brief starts the build. Every artifact generates the next one. ## Further Reading - [The Index](/case-studies/the-index): the 256-row evidence base the tools draw from. - [Strength of Evidence](/case-studies/strength-of-evidence): which published numbers hold up, and how to quote them responsibly. - [Reference Index](/reference): every published Anthropic source, dated.