← All projects

CLEW Agent

Support drafts a human approves

CLEW Agent drafts a reply as an internal note for every incoming Zendesk ticket for the snowboard brand CLEW. A human reviews and sends it, the agent never sends anything itself.

I built the pipeline, knowledge base and guardrails, and took the system into production.

Status
Live since August 2026
Platform
Zendesk, Shopify widget
Role
Architecture, development, operations

Case study by Fabian Bitzer 5 min read

CLEW Agent is an AI system for the snowboard brand CLEW that drafts a reply to every incoming Zendesk ticket, which a human then reviews and sends. It looks up orders across two Shopify stores, pulls knowledge from the helpdesk and support macros, and follows 24 fixed guardrails. I built the architecture, pipeline and operations.

Why CLEW needed a new support agent

CLEW is a snowboard brand running two Shopify stores, one for Europe and one for the US and Canada, plus its own helpdesk. Support tickets about bindings, spare parts, shipping and warranty all arrive through Zendesk. A previous system had no cap on model calls and no fixed rules for its answers. Many rules in CLEW Agent go back to failures documented there.

I built the agent end to end, from architecture to operations. The core decision came first: the agent never sends anything itself. For every incoming ticket it writes a reply draft as an internal note in Zendesk, and a human reviews it before anything goes out. That is enforced in code, not just policy: the Zendesk client has no method for a public reply at all.

How a draft gets written

The Zendesk webhook carries only the ticket number, the agent fetches everything else itself through the API. Seven fixed steps run for every ticket, and at most three of them touch a language model.

Step What happens
Noise filter spam and auto-replies dropped, no model call
Fetch context Zendesk thread and both Shopify stores, in parallel
Classify one model call, returns category, language, part, order number
Decide warranty, escalation, reseller routing, code only, no model
Write the draft second model call, grounded in helpdesk, catalog and macros
Guardrails 24 fixed rules check the text, in code
Repair or template one more attempt, otherwise a hand-written fallback

The result is an internal note plus pre-filled ticket fields: order number, name, address, tracking number, topic. Only empty fields get filled, a colleague’s own input is never overwritten.

  1. Noise filter
  2. Fetch context
  3. Classify
  4. Decide
  5. Write draft
  6. 24 guardrails
  7. Fix or template

internal note

Never sends on its ownthe client has no public reply method

max. 3 model calls

Seven fixed steps per ticket. Only the three marked ones call a language model, and a single wrapper function enforces that limit. The result is an internal note in Zendesk: a person reads it and sends the reply.

Knowledge from the helpdesk, the catalog and real replies

The knowledge base lives in Postgres with pgvector, split into two namespaces: public knowledge, meaning helpdesk articles and the product catalog, and internal support macros that must never reach a customer-facing context. The helpdesk is crawled rather than read through the official API, because that API returns an empty HTML shell for part of the articles, including the repair guides for straps and bindings.

Each draft runs three separate searches instead of one, with reserved slots for macros, helpdesk articles and products. Before that change, a warranty question competed with the product catalog for the same slots and lost, simply because there are far more product chunks than articles.

An embedding model chosen by measurement

Six candidates were benchmarked against twenty real questions from CLEW’s own helpdesk, mostly German questions against English articles. The winner was text-embedding-3-large, truncated to 1024 dimensions. The model used before came in last.

Guardrails: what the agent never does

24 hard rules check every draft, each with its own test, several taken directly from a documented failure of the predecessor. The agent never promises a free replacement or refund, never announces a handover, never signs with a colleague’s name, and never deviates from a fixed fact like the return window. If a check fails, the model gets exactly one repair attempt, after which a hand-written template takes over.

A chat widget for the store, too

The same architecture also powers a public chat widget for the Shopify storefront, embedded with a single script tag. Its look is not a separate design system but read directly from the shop’s own theme variables: square corners, no shadows, one typeface, no third-party CDN.

CLEW's digital assistant in German: a greeting, the question where to find spare parts for the Independence binding and the answer with a link to the spare parts page.
The chat widget in CLEW's look (German): square corners, no shadows, one typeface. Here it answers a question about spare parts with a link.

An order only becomes visible once an order number and a shipping address are confirmed together against the shop, checked in code before a prompt is even built. Without that check, the order simply does not exist for the model, no matter how the chat message is phrased. If the bot cannot help, it opens a ticket as an internal note rather than replying publicly, so no automatic email goes to an address that might be mistyped.

Architecture and stack

TypeScript on Node 22 in strip-only mode, so no build step. Hono as the web server, the Vercel AI SDK with OpenRouter for model calls, Postgres with pgvector through Mastra, kept deliberately thin: Mastra owns the database connection and memory, not a single pipeline decision. It runs on Railway, deployed from GitHub, with the main branch protected so every change goes through a pull request.

The model profile switches with a single variable, no code change required: one profile for a frontier model, one restricted to EU-based providers for stricter data protection needs, one for a self-hosted model.

Measured, not assumed

Against a golden set of 2,142 real, anonymized cases, the agent reaches 99.7 percent correctly identified noise, 84.2 percent of escalation decisions matching a human’s, 98.3 percent correct reply language, and zero drafts with a guardrail violation. The noise filter alone reaches 100 percent precision.

What I learned building it

A mock only proves that code and mock agree with each other. A hand-written test assumed order numbers carried a leading hash; the real shop returns them without one. Hundreds of passing tests hid that until a real case surfaced it. Fixtures now come from real, anonymized responses with their own redaction step.

An enforced limit beats a documented rule. The old three-call limit used to live in a note somewhere. Now it sits inside a single function wrapping every model call, so a call over budget never reaches the provider at all.

A noise filter pays off most when it runs first. It costs no model call, and alone would have prevented hundreds of unnecessary runs on the predecessor system.

Status and access

CLEW Agent has been running in production since late August 2026, drafting replies to real Zendesk tickets for CLEW. The system is an internal tool with no public interface of its own. It only shows indirectly, in the replies of the support team.

Questions about the project, or something similar for your business? Write to me at kontakt@bitzer-fabian.de.

Next project

CLEW NFC

An NFC chip in every snowboard binding: customers tap to register, and CLEW gets its direct customer contact back.