← All projects

MMK Chatbot

Study questions answered with sources

MMK Chatbot answers study questions for Digital Media (MMK) at DHBW Mannheim: deadlines, exams, phase plans and procedures. It reads the official documents, answers in full sentences and links the document the answer comes from.

I built it as a student assistant for the program and as the practical part of my bachelor's thesis on privacy-friendly RAG at universities: hybrid search over the knowledge base, one login per student with the anonymous S-address, saved chats and feedback right in the chat.

Status
Live
For
Digital Media (MMK) students, DHBW Mannheim
Knowledge
588 documents and pages, 5,317 sections
Stack
Next.js, FastAPI, Qdrant, PocketBase
Role
Design, engineering, operations

Case study by Fabian Bitzer 6 min read

The MMK Chatbot answers questions from Digital Media students at DHBW Mannheim using the program's official documents, and links the source for every answer. Under the hood it is a RAG system with hybrid search across 588 documents and pages. I built it as the practical part of my bachelor's thesis and as a student assistant project for the program.

Why a chatbot for a study program?

Students of Digital Media with a focus on media management and communication (MMK) at DHBW Mannheim have organisational questions all the time. By when does a sick note have to be submitted? How do I withdraw from an exam? How do I apply for leave? What are the formal rules for the bachelor’s thesis? When is the next practical phase? The answers are spread across study and examination regulations, phase plans, guidelines and dozens of web pages. Whoever cannot find them writes to the program office.

The MMK Chatbot takes that search off students’ hands. They ask in their own words and get an answer in full sentences, with a link to the document it comes from. Two things it deliberately does not do: it gives no binding legal advice and refers students to the program office in those cases, and it is not a conversation partner for personal topics.

Home page of mmk-chatbot.com in German: the heading MMK Chatbot, the tagline Your studies, answered with a source, and a source card from the Digital Media module handbook.
The information page at mmk-chatbot.com (German). The chatbot itself runs at app.mmk-chatbot.com, and students sign in with their S-address.

What is a RAG system?

RAG stands for retrieval-augmented generation. A RAG system answers a question in two steps. First it searches its own knowledge base for the passages that match the question. Then a language model writes the answer based only on those passages and names its sources. That keeps answers current and verifiable, and the model does not have to make anything up from memory. For a study program full of regulations and deadlines, that is the decisive advantage over a general-purpose AI.

  1. question

    When do I have to be at the company?

    glossary at the company phase planning
  2. hybrid search

    Semantic, embeddings
    Lexical, BM25

    RRFone merged ranking

  3. context

    sections per answer mode

  4. answer, streamed

    Phase planning MMK23, PDF

From question to cited answer. The glossary translates everyday language, two searches look in parallel, reciprocal rank fusion merges the lists, and only the selected sections reach the language model.

How the MMK Chatbot works

The knowledge base

It starts with a raw corpus of several thousand crawled DHBW web pages and PDFs. Using an explicit per-file decision list plus rules, I filtered it down to a focused set: bachelor’s level, business faculty, MMK track. I deliberately excluded the master’s program that shares the same abbreviation, because it led to mixed and therefore wrong answers.

Today the knowledge base covers 588 documents and pages, split into 5,317 sections. Of 575 source links, 571 are verified as reachable, so practically every citation leads to the original.

raw crawl

several thousand pages and PDFs

filter, file by file

  • ✓Bachelor's level
  • ✓Business faculty
  • ✓MMK track
  • ✕Master's program MMK

knowledge base

588documents and pages

5,317sections

571 of 575source links verified reachable

The knowledge base is deliberately narrow: from several thousand crawled pages and PDFs, only bachelor's level, business faculty and the MMK track remain. The master's program with the same abbreviation is excluded.

Preparing the texts

Documents are split into sections of 320 words with a 60-word overlap. Each section also gets a generated context sentence and five to ten keywords. On top of that, a glossary translates students’ everyday language into DHBW vocabulary. A question like “When do I have to be at the company?” therefore finds the phase planning, even though the word phase planning never appears in the question.

  1. section 1 · 320 words

    context sentence

  2. section 2 · 320 words

    context sentence

  3. section 3 · 320 words

    context sentence

Schematic: documents are split into sections of 320 words with a 60-word overlap, so text at a boundary appears in both neighbouring sections. Each section gets a generated context sentence and five to ten keywords.

Search runs in Qdrant and combines two methods: a semantic search over embeddings that finds meaning, and a lexical BM25 search that finds exact technical terms. Both result lists are merged with reciprocal rank fusion. Depending on the answer mode, 6, 12 or 20 sections go to the language model.

The answer

The language model runs through a gateway in EU data centres. The answer is streamed word by word into the interface with server-sent events, and the sources appear once the answer is complete. The backend itself is built with FastAPI and stores nothing.

Interface, sign-in and feedback

The interface is a custom Next.js application. I evaluated existing open-source chat interfaces and then chose a custom one, tailored exactly to source display, sign-in and feedback.

Every student gets their own account based on their anonymous DHBW S-address. There is no self-registration, and a new password is mandatory on first sign-in. Accounts and saved chat history live in PocketBase. Every answer can be rated with a thumbs up or down. Before a conversation is stored as feedback, the browser already removes email addresses, IBANs, phone numbers and long digit sequences.

How well does the system find the right passage?

A RAG system is only as good as its search. If the search misses the right passage, even the best language model cannot answer correctly. So I measure the search with my own gold set of 40 hand-written questions in three groups: colloquially paraphrased questions, literal technical terms and questions in English.

The key metric is Recall@12: the share of questions for which the correct document is among the first twelve results. After the search overhaul in July 2026, the numbers looked like this:

Metric Before After
Recall@12, paraphrased questions 0.45 0.50
Recall@12, literal technical terms 0.80 0.933
Recall@12, overall 0.625 0.70
MRR 0.406 0.455
nDCG@12 0.613 0.718

After moving to new models for answers and embeddings in September 2026, the same gold set reached a Recall@12 of 0.825 and an MRR of 0.516, with a recall of 0.65 on paraphrased questions. The previous configuration was not re-measured in that run, so the comparison shows a trend rather than an exact side-by-side.

Recall@12, paraphrased questions
0.45 0.50 Sept. 0.65
Recall@12, literal technical terms
0.80 0.933
Recall@12, overall
0.625 0.70 Sept. 0.825
MRR
0.406 0.455 Sept. 0.516
nDCG@12
0.613 0.718
Search quality on my gold set of 40 questions. Grey: before the overhaul in July 2026, black: after it. The outlined marks come from a separate run after the model switch in September 2026 and show a trend, not an exact side-by-side.

What did not work

The most instructive finding of the project: a well-known recommendation turned out to be wrong for this corpus. With Contextual Retrieval, Anthropic describes a method in which every section gets a generated context sentence that goes into both the semantic and the lexical index. Anthropic reports significantly fewer failed retrievals.

I measured this directly in the MMK Chatbot. In the semantic index, the context sentence made recall on paraphrased questions worse, from 0.45 to 0.40, while it helped with technical terms and English questions. The reason lies in the material: university documents consist to a large extent of nearly identical administrative texts. A context sentence per section ends up describing the whole document every time and makes sections more alike rather than more distinguishable. The fix: the context sentence stays in the lexical index only.

paraphrased questions, Recall@12

0.45without context sentence 0.40with context sentence

where the context sentence goes

  • Semantic index off
  • Lexical index, BM25 on

helps: technical terms helps: English questions

Measured, not assumed: in the semantic index the context sentence lowered recall on paraphrased questions from 0.45 to 0.40. It stays in the lexical index only, where it helps with technical terms and English questions.

Other ideas did not hold up either. Making file names searchable did not improve recall. A JSON format for the generated keywords broke on unescaped quotation marks. And a token limit set too tight silently cut off models with a reasoning phase. The general lesson: recommendations are hypotheses, and decisions are made with measurements on your own corpus.

Privacy without dedicated hardware

The prototype for my bachelor’s thesis was hybrid: a router checked every question for personal data with Microsoft Presidio and sent sensitive questions to a local model running on Ollama, while everything else went to a cloud model. For production I removed that path on purpose. The knowledge base content is mostly public, affordable models in EU data centres address the privacy question sufficiently, and dedicated hardware would not have saved any money.

The chatbot runs on a server managed with Coolify, which keeps deployments and continuous operation simple. There is intentionally no admin interface for accounts: new cohorts are imported by script from a list of S-addresses. Anything more would be oversized for the purpose.

Connection to my bachelor’s thesis

The MMK Chatbot is the practical artefact of my bachelor’s thesis at DHBW Mannheim on a hybrid, privacy-compliant RAG architecture for public universities. The differences between prototype and production, above all dropping the local model, are a result in their own right: they show which architecture actually works for public content in everyday university life.

Status and access

The application is live at app.mmk-chatbot.com, with an information page at mmk-chatbot.com (German). It is available to students of the MMK track with their S-address. Are you working on a similar system for a university or an organisation with lots of documents? Write to me at kontakt@bitzer-fabian.de.

Next project

Sharehood

A neighborhood app for borrowing tools, kitchen appliances and more, map-based and hyperlocal, built as a solo project.