Resources

Playbooks, benchmarks, and results from the queue.

How teams put AI agents on their highest-volume work without giving up control. Grounded in real deployments — with the numbers, the flows, and widgets you can play with.

Case Study··8 min

How Cenoa cut business onboarding from 2 weeks to 2 days

A digital bank for businesses turned a 2-week KYB backlog into a 2-day, day-one-revenue flow — with a compliance reviewer on every approval. The numbers, the flow, and what carried over.

Case Study··8 min

How Getir ran 200K product onboardings a month with 5 people, not 25

6,000 images a day, a 25-person QA bottleneck, and GMV lost every hour a restaurant stayed offline. Here is the catalog-review flow that took time-per-task from 90 seconds to 1.5.

Benchmark··7 min

The cost of slow onboarding: 70% of firms lose a client over it

70% of financial institutions have lost a client because business onboarding was too slow (Fenergo, 2024). We break down where the days go and what each one costs.

Case Study··7 min

SAP order management: from 15 minutes to 30 seconds per order

B2B order entry is copy-paste between email, PDF, and SAP. Here is how an agent handles the routine orders end-to-end and hands a human the exceptions, cutting 15 minutes to 30 seconds.

Benchmark··7 min

25 days to first sale: the marketplace activation benchmark

Across 65,000 sellers, average time from signup to first sale on enterprise marketplaces is 25 days (Mirakl, 2023). Catalog review is the hidden tax. Here is the math.

Case Study··7 min

Intelligent document processing in logistics: same-day, 98% accuracy

BOLs, customs forms, and invoices arrive as PDFs and photos. Here is the document-processing flow that hit same-day turnaround at 98% accuracy — with humans on the exceptions.

Benchmark··6 min

$250 per exception PO: the real cost of invoice mismatches

A single purchase order that doesn't match its invoice costs $250 fully loaded to resolve (APQC). Multiply by your exception rate. Here's how to bring it down.

Benchmark··7 min

Getting past the 78% wall: the accuracy curve nobody shows you

Most agent pilots flat-line near 78% accuracy. The fix isn't a bigger model — it's supervisor corrections turned into training signal. Watch the curve climb, week by week.

Benchmark··8 min

Build in-house vs platform: 6–12 months vs 3 weeks

The build-vs-buy spreadsheet always undercounts the same things: eval infra, audit trails, and the engineering ticket behind every workflow change. A grounded comparison.

Industry··7 min

Clearing peak-season support volume without hiring for it

A flash sale triples the queue overnight. WISMO, refunds, and fraud holds pile up faster than temps can ramp. Here's the human-in-the-loop pattern that absorbs the spike.

Blog··7 min

It's not an automation problem. It's a deployment problem.

You bought a workflow tool. Your team configured it for months. The AI made errors nobody could explain, and the project got shelved. That's a deployment problem — here's the fix.

Industry··7 min

Catalog onboarding at scale: review thousands of SKUs a day

Every new SKU needs classification, copyright checks, and quality review before it can sell. Here's how to run catalog onboarding at scale with a reviewer on the edge cases.

Blog··7 min

An operating model, not a tool you buy

Workflow builders need rules. Internal tools need a developer. RPA breaks in production. Full agents leave no one accountable. Here's the operating model that avoids all four traps.

Industry··8 min

KYC/KYB onboarding without a compliance backlog

Document checks, sanctions screening, and risk scoring are the reason activation stalls. Here's how to clear the routine cases and put a reviewer on every flag.

Blog··7 min

The middle path: why augmented ops is the hardest thing to build

100% manual doesn't scale and burns out operators. Full automation is easy to break and impossible to govern. The middle — AI runs the flow, a human owns the last call — is the hard part.

Industry··7 min

Transaction dispute resolution at scale, with an audit trail

Chargebacks and disputes are high-volume, high-stakes, and deadline-driven. Here's how agents assemble the evidence pack and route the judgment calls to a human in seconds.

Blog··6 min

How the same team processes 10x the volume

The 10x isn't a bigger headcount or a faster tool — it's a flow where the agent moves optimistically through every case and the operator approves in one click. Here's how it works.

Industry··7 min

Shipment exception triage: catching SLA risk before the penalty

The 10% of shipments that go sideways eat most of the ops day. Here's how to triage customs holds, delays, and damage claims across the TMS, email, and carrier portals.

Blog··5 min

Why human-in-the-loop AI agents win in production

Autonomous agents stall at the trust bar. Human-in-the-loop ships. Here is the architecture that gets AI past the pilot stage and into real ops.

Blog··6 min

No rip-and-replace: put agents where the work already lives

SAP, Slack, HubSpot, your warehouse — the systems aren't the problem, the handoffs between them are. Here's how to deploy agents into your existing stack without a migration.

Industry··7 min

Intelligent document processing for freight paperwork

BOLs, customs declarations, proofs of delivery — the paperwork is the job. Here's how to extract, validate, and reconcile it with confidence routing and a clean audit trail.

Blog··6 min

What “production-grade” actually means for AI agents

A working definition of production-grade for AI agents: confidence thresholds, fallbacks, rollback, audit, and an eval harness that survives a model swap.

Industry··7 min

Supplier onboarding: from weeks to days before the first PO

Certifications, document checks, and compliance review across disconnected systems keep a new vendor waiting weeks. Here's how to compress that to days without dropping controls.

Benchmark··7 min

The eval harness gap: why agent pilots stall at 78% accuracy

Most AI pilots flat-line around 78% accuracy. The bottleneck is rarely the model. It is the missing eval harness. Here is how to build one.

Industry··7 min

Quality inspection triage before defects ship downstream

Inspection reports, defect photos, and return-material requests outpace the QA team. Here's how to triage them fast and put an engineer on the critical calls.

Blog··6 min

How to scope a 30-day AI agent pilot for ops teams

A practical 30-day playbook to take an ops workflow from messy spreadsheet to production agent without lighting compliance on fire.

Industry··7 min

Real-time payout and dispute ops for live events

When an event goes live, payouts, disputes, and support spike in minutes. Here's how to clear the routine cases in real time and route the risky ones to a human.

Blog··5 min

Confidence thresholds, explained: routing decisions to humans

Confidence thresholds are the steering wheel of a human-in-the-loop agent. Here is how to pick them, tune them, and avoid the common traps.

Industry··7 min

Content moderation at scale, with receipts

Moderation and creator disputes need speed and a defensible record. Here's how agents handle the clear cases in real time and hand humans the judgment calls, logged.

Blog··6 min

Audit trails for AI agents: a compliance primer

What an auditable AI agent actually logs, why it matters for SOX, GDPR, and HIPAA reviewers, and what to ask your vendor before you sign.

Industry··7 min

First-pass claims adjudication in minutes, not days

FNOL, document intake, and policy checks are repetitive until they aren't. Here's how to adjudicate the clean claims end-to-end and hand adjusters the ambiguous ones, evidence-ready.

Blog··5 min

Why supervisor corrections are the moat

Compounding accuracy is the only loop we have seen consistently beat the 78% wall. The fuel is supervisor corrections, structured into training signal.

Blog··6 min

Choosing your first ops workflow for AI automation

A scoring rubric to pick the workflow with the best chance of clearing a 30-day pilot, without betting the quarter on a moonshot.

Industry··7 min

Tier-1 ticket triage with HITL agents: what to expect

Volume, accuracy, escalation rates, supervisor load. A real-world look at running tier-1 ticket triage on a human-in-the-loop agent.

Industry··6 min

Vendor onboarding automation: where the bottlenecks actually live

Most vendor onboarding teams blame procurement. The real time-sinks are paperwork parsing, sanctions screening, and ERP data entry. Here is the fix.

Industry··7 min

First-pass claims review: patterns, anti-patterns, and pitfalls

What separates a claims review agent that actually ships from one that lives forever in a sandbox. Patterns, anti-patterns, and the pitfalls we have hit.

Blog··5 min

Lead enrichment with AI: scores you can defend in QBR

An ICP score nobody trusts is worse than no score. Here is how to build lead enrichment that reps lean on and that survives a quarterly review.

Benchmark··7 min

Self-built vs platform: the real cost of in-house agent infra

We built it twice. Here is what we underestimated, what surprised us, and the question to ask before you spin up the build-vs-buy spreadsheet.

Blog··5 min

Per-seat vs volume pricing for AI agents (and why we chose volume)

Per-seat pricing punishes the workflows AI agents are best at: high volume, low headcount. Here is the math behind volume-based pricing.

Blog··6 min

When NOT to use an AI agent: a checklist for ops leaders

Not every workflow is an agent workflow. A six-question checklist to help ops leaders decide when to ship a HITL agent and when to walk away.