Docs
How Legacy Loom works
Everything you need to run it, deploy it, and understand why it answers the way it does.
Overview
Legacy Loom keeps one person's voice notes and makes them useful. You add recordings, it writes them down with timestamps, files each one (what kind of memory it is, who is in it, when it happened, the lines worth keeping), and lets you ask questions. Answers come only from the recordings and every claim links to the second it was said.
It is built so the core runs on open models. Gemma does the reading and answering, Whisper does the listening, an open embedding model does the searching. On a laptop all three run locally and the voice notes never leave it.
Quickstart
You need Python 3.11 or newer, Ollama and a free MongoDB Atlas cluster.
ollama pull gemma3:4b
python -m venv .venv
.venv\Scripts\activate # macOS or Linux: source .venv/bin/activate
pip install -r requirements-local.txt
copy .env.example .env # fill in MONGODB_URI and any keys you have
python -m scripts.setup_atlas # creates the vector and text search indexes
uvicorn app.main:app --port 8000
Open http://127.0.0.1:8000/app, go to Settings, enter whose voice it is, then add a recording. The first one takes a minute or two on a CPU while Whisper and Gemma load.
From voice note to memory
- Store. The audio goes into GridFS so it can be streamed back later with byte ranges, which is what lets the player jump straight to a cited second.
- Listen. faster-whisper transcribes on the CPU with word timings. When it is not installed (as on a serverless host) ElevenLabs Scribe does it instead. Both return the same shape: segments, words, language, duration.
- File. Gemma reads the transcript and returns JSON: a title, the kind of memory (story, recipe, advice, song), a summary, people, places, a year when one is said, topics, up to three exact quotes, and a recipe card when a dish is explained. If a small model marks a note as a recipe but leaves the card empty, a second focused pass asks for just the card.
- Index. The transcript is grouped into passages of about 45 seconds that keep their timestamps. Each passage plus one English overview passage is embedded with
BAAI/bge-small-en-v1.5and stored with its vector. - Log. A session row records when the recording happened, who asked, the topic, how long they spoke and how fast. The call planner learns from these rows.
How answers are made
Each question goes through a small agent loop:
- Route. Gemma picks one or two tools:
search_memories,find_recipe,timelineorplan_calls. - Retrieve. Search runs Atlas Vector Search and Atlas Search (keyword, with typo tolerance) side by side and merges them with reciprocal rank fusion, so a passage that ranks well in both rises to the top.
- Answer. Gemma writes the answer from the numbered material only and cites it like
[2]. If nothing in the archive answers the question, it says so and suggests a question to ask next time. - Check. Every quoted span in the answer is compared with the real transcripts. A sentence quoting words that were never said is removed before the answer is final, and the interface says so.
The call planner
Rate recordings from one to five in the Archive. Four or five counts as a keeper. The planner gives TabPFN a small table with one row per session: hour, weekday, who asked and the topic. It learns which combinations tend to produce keepers, then scores every hour and topic for the next seven days. A second TabPFN model estimates how many minutes they will talk.
It needs at least 8 rated sessions, with some keepers and some that were not. Until then it says how many more it needs. You can import an existing call log as CSV:
date,time,duration_min,asked_by,topic,rating
2026-09-12,19:30,14,Ada,jollof,5
Configuration
| Variable | What it does |
|---|---|
LLM_PROVIDER | ollama runs Gemma on this machine, google uses Google AI Studio. |
OLLAMA_MODEL | Local model, default gemma3:4b. |
GOOGLE_API_KEY, GOOGLE_MODEL | Hosted Gemma. Default model gemma-4-26b-a4b-it. |
MONGODB_URI, MONGODB_DB | Atlas connection and database name. |
TRANSCRIBER | auto, local or elevenlabs. |
WHISPER_MODEL, WHISPER_LANGUAGE | Whisper size (default small) and an optional forced language. |
ELEVENLABS_API_KEY, ELEVENLABS_VOICE_ID | Read aloud and Scribe transcription. |
TABPFN_TOKEN | Prior Labs API token for the planner. |
APP_PASSCODE | When set, adding, rating and deleting need it. Set it before sharing a link. |
PUBLIC_BASE_URL | Used for the QR codes in the recipe book. |
Deploying
Vercel
The site and the API deploy from vercel.json. Pages are served from the CDN and the API runs as a Python function. Two things are different on serverless and the code handles both:
- A function is frozen once it replies, so an upload is transcribed and filed inside the upload request.
- The TabPFN client needs about 400 MB of libraries, more than fits beside the app in one function, so the planner deploys as its own small project (
tabpfn_api/) and the main site proxies/api/plannerto it.
Render
render.yaml describes a single web service. It runs the planner in process and processes uploads in the background.
API reference
| Method and path | What it does |
|---|---|
GET /api/status | Which services are configured and reachable. |
GET /api/memos | All memos, newest first. |
POST /api/memos | Upload a recording (multipart: file, recorded_at, asked_by, prompt_topic). |
GET /api/memos/{id} | One memo with transcript segments. |
PATCH /api/memos/{id} | Rate, retitle or correct the year. |
GET /api/memos/{id}/audio | The recording, with byte range support. |
POST /api/ask | Ask a question. Streams server sent events: plan, tool, sources, token, removed_quotes, done. |
POST /api/tts | Read text aloud (MP3). |
GET /api/planner | The TabPFN forecast. |
POST /api/sessions/import | Import a call log CSV. |
GET /book | The printable recipe book. |
Every call can carry an x-space header: main for the archive, try for the open sandbox. In main, anything that changes data needs the x-family-key header when a passcode is set. GET /api/access tells you whether you can write.
Privacy
On a laptop with Ollama and Whisper, transcription, filing and answering happen on the machine. Only the database is remote, and Atlas can be swapped for a local MongoDB if you want everything offline (vector search then needs Atlas or a local Atlas deployment).
The hosted version sends audio to ElevenLabs for transcription and transcripts to Google AI Studio for Gemma. That trade is made for a demo anyone can open. For a real family archive, run it locally.
Voice cloning is never turned on by default. Only do it with the person's clear permission.
Known limits
- Uploads on Vercel are capped at 4 MB per file, which is still over 30 minutes of a WhatsApp voice note.
- Whisper handles English, Yoruba and Hausa unevenly and does not cover Igbo. Use ElevenLabs Scribe for languages it struggles with.
- The local 4B model is slower and less careful than the hosted 26B. The quote check catches invented quotes from either.
- The planner is only as good as the ratings. With fewer than 8 it does not guess.