Karla is the AI chatbot on karlsruhe-erleben.de, the official tourism portal for the city of Karlsruhe. It answers questions about events, sights and restaurants live from the city's own event database and the website's content. I am responsible for the technical implementation: the chatbot, its knowledge base, quality assurance and the monthly client report.
A travel guide for an entire city
Karlsruhe Marketing und Event (KME) and Karlsruhe Tourismus GmbH (KTG) run karlsruhe-erleben.de, the city’s official tourism portal. Visitors ask it the same questions they would ask at any tourist office: what is happening today, where can I get a good meal, how do I best get into the city. The answers live scattered across an events database, hundreds of place descriptions and dozens of pages about arrival, districts and sights.
Karla is the chatbot that answers those questions directly on the website, in German, English and French, around the clock. The project is led by Prof. Dr. Gerald Lembke, who is responsible for the content and quality side as project lead. I took on the technical implementation: the whole chatbot architecture, the connection to the data sources, quality assurance and the ongoing reporting to the client.
How Karla answers a question
Karla is not a single language model but an agent with several tools that decides, question by question, where the answer should come from:
- Event questions (“what’s on this weekend?”) go to a dedicated tool that queries the KTG’s Toubiz API live, the same database that feeds the event calendar on the website itself. A second tool always supplies the current date first, because the language model has no built-in sense of today’s date or time, and without that step it would reliably get “today” and “tomorrow” wrong.
- Questions about places (restaurants, sights, public restrooms, opening hours) run through a second tool against the same Toubiz database, this time its places and articles endpoint.
- General questions about Karlsruhe (getting there, history, districts) are answered from Karla’s own knowledge base, built from the website’s editorial content.
This split is the core of the architecture: wherever a structured, current source exists, Karla uses it live instead of relying on a text passage that might be outdated. The system prompt makes this an explicit rule: Karla only answers based on tool results and never invents opening hours, prices or addresses.
question
What's on this weekend?
tooltoday's date fetched
- EventsToubiz API, live
- PlacesToubiz API, places
- General questionsKnowledge base, Qdrant
Karla
Event calendar
DEENFR
Architecture
| Component | Role |
|---|---|
| Flowise | Orchestrates the agent, its tools and the conversation |
| Qdrant | Vector database for the website content |
| Language model via Requesty | EU-hosted gateway, currently a Gemini Flash model |
| Toubiz API | Live data source for events and places |
| Custom ingest pipeline | Crawls, chunks and updates the knowledge base |
The chatbot itself runs in Flowise, an open-source tool for AI agents, self-hosted on its own server. The running version is pinned deliberately, and every change is tested against a throwaway copy first, never directly against the live bot.
For the knowledge base I run a custom ingest pipeline that crawls the website content, splits it into sections and embeds it into Qdrant. It started out as an automated nightly job, but now runs on demand through a small internal tool: an editor enters the page that changed, the pipeline re-reads it and replaces only the affected sections. That turned out to be more robust than a full nightly run across thousands of pages, where a single page’s errors get lost in the noise.
Quality assurance: tested against real questions, not gut feeling
Before a new language model or a new prompt goes live, it runs against a fixed test matrix of several dozen real user questions, from event questions to arrival information to questions in English and French. Every answer is checked automatically against fixed criteria, such as whether an expected source link is present or whether forbidden formatting shows up. On that basis I systematically benchmarked several language models against each other before one of them went into production.
On top of that sits a deny list with several hundred entries that keeps Karla from engaging with manipulation attempts or inappropriate requests, staying factual or pointing to the tourist office instead. I check these rules in live operation too, not only before a rollout.
Monthly reporting for the client
Beyond the chatbot itself, I built a dedicated reporting system that gives KME a monthly evaluation report. A script first computes every metric from the conversation export, then a group of AI agents categorizes every single conversation of the month by topic, reviews critical requests in detail and rates answer quality. Both feed into a self-contained HTML dashboard that goes straight to the client team by email.
Building the report deliberately fails if that qualitative review is missing, so a report never goes out as bare numbers without real interpretation.
What I learned
The chatbot itself was not what took the most time, the data connection was. A public API being documented does not mean the documentation is complete. Several parameters of the Toubiz API behaved differently from what the docs said, and the difference only showed up when testing against real responses, not while reading the docs. Since then, my rule for any third-party API is to verify against the actual response, not the description.
The second lesson is about state versus call parameters. A setting that only exists as a parameter of a single call, with nowhere to persist it, silently reverts the next time that call runs without the parameter, even if it had been set correctly before. That taught me to treat configuration as stored state by default, not as something passed in on every call.
Status
Karla runs in production on karlsruhe-erleben.de and is under continuous development, most recently with a language model switch in September 2026. Are you building something similar, a chatbot wired into existing data sources, or reporting for recurring evaluations? Write to me at kontakt@bitzer-fabian.de.