Docs

How Legacy Loom works

Everything you need to run it, deploy it, and understand why it answers the way it does.

Overview

Legacy Loom keeps one person's voice notes and makes them useful. You add recordings, it writes them down with timestamps, files each one (what kind of memory it is, who is in it, when it happened, the lines worth keeping), and lets you ask questions. Answers come only from the recordings and every claim links to the second it was said.

It is built so the core runs on open models. Gemma does the reading and answering, Whisper does the listening, an open embedding model does the searching. On a laptop all three run locally and the voice notes never leave it.

Where things live. Memos, searchable passages, sessions and the audio itself are stored in MongoDB Atlas. The code is on GitHub under the MIT license.

Quickstart

You need Python 3.11 or newer, Ollama and a free MongoDB Atlas cluster.

ollama pull gemma3:4b
python -m venv .venv
.venv\Scripts\activate            # macOS or Linux: source .venv/bin/activate
pip install -r requirements-local.txt
copy .env.example .env            # fill in MONGODB_URI and any keys you have
python -m scripts.setup_atlas     # creates the vector and text search indexes
uvicorn app.main:app --port 8000

Open http://127.0.0.1:8000/app, go to Settings, enter whose voice it is, then add a recording. The first one takes a minute or two on a CPU while Whisper and Gemma load.

From voice note to memory

  1. Store. The audio goes into GridFS so it can be streamed back later with byte ranges, which is what lets the player jump straight to a cited second.
  2. Listen. faster-whisper transcribes on the CPU with word timings. When it is not installed (as on a serverless host) ElevenLabs Scribe does it instead. Both return the same shape: segments, words, language, duration.
  3. File. Gemma reads the transcript and returns JSON: a title, the kind of memory (story, recipe, advice, song), a summary, people, places, a year when one is said, topics, up to three exact quotes, and a recipe card when a dish is explained. If a small model marks a note as a recipe but leaves the card empty, a second focused pass asks for just the card.
  4. Index. The transcript is grouped into passages of about 45 seconds that keep their timestamps. Each passage plus one English overview passage is embedded with BAAI/bge-small-en-v1.5 and stored with its vector.
  5. Log. A session row records when the recording happened, who asked, the topic, how long they spoke and how fast. The call planner learns from these rows.

How answers are made

Each question goes through a small agent loop:

  1. Route. Gemma picks one or two tools: search_memories, find_recipe, timeline or plan_calls.
  2. Retrieve. Search runs Atlas Vector Search and Atlas Search (keyword, with typo tolerance) side by side and merges them with reciprocal rank fusion, so a passage that ranks well in both rises to the top.
  3. Answer. Gemma writes the answer from the numbered material only and cites it like [2]. If nothing in the archive answers the question, it says so and suggests a question to ask next time.
  4. Check. Every quoted span in the answer is compared with the real transcripts. A sentence quoting words that were never said is removed before the answer is final, and the interface says so.
The quote check exists because a small model, asked to be warm, once wrote "She always said, it's a good, warming soup." Nobody said that. An archive of someone's voice cannot make that mistake, so it is checked, not trusted.

The call planner

Rate recordings from one to five in the Archive. Four or five counts as a keeper. The planner gives TabPFN a small table with one row per session: hour, weekday, who asked and the topic. It learns which combinations tend to produce keepers, then scores every hour and topic for the next seven days. A second TabPFN model estimates how many minutes they will talk.

It needs at least 8 rated sessions, with some keepers and some that were not. Until then it says how many more it needs. You can import an existing call log as CSV:

date,time,duration_min,asked_by,topic,rating
2026-09-12,19:30,14,Ada,jollof,5

Configuration

VariableWhat it does
LLM_PROVIDERollama runs Gemma on this machine, google uses Google AI Studio.
OLLAMA_MODELLocal model, default gemma3:4b.
GOOGLE_API_KEY, GOOGLE_MODELHosted Gemma. Default model gemma-4-26b-a4b-it.
MONGODB_URI, MONGODB_DBAtlas connection and database name.
TRANSCRIBERauto, local or elevenlabs.
WHISPER_MODEL, WHISPER_LANGUAGEWhisper size (default small) and an optional forced language.
ELEVENLABS_API_KEY, ELEVENLABS_VOICE_IDRead aloud and Scribe transcription.
TABPFN_TOKENPrior Labs API token for the planner.
APP_PASSCODEWhen set, adding, rating and deleting need it. Set it before sharing a link.
PUBLIC_BASE_URLUsed for the QR codes in the recipe book.

Deploying

Vercel

The site and the API deploy from vercel.json. Pages are served from the CDN and the API runs as a Python function. Two things are different on serverless and the code handles both:

  • A function is frozen once it replies, so an upload is transcribed and filed inside the upload request.
  • The TabPFN client needs about 400 MB of libraries, more than fits beside the app in one function, so the planner deploys as its own small project (tabpfn_api/) and the main site proxies /api/planner to it.

Render

render.yaml describes a single web service. It runs the planner in process and processes uploads in the background.

API reference

Method and pathWhat it does
GET /api/statusWhich services are configured and reachable.
GET /api/memosAll memos, newest first.
POST /api/memosUpload a recording (multipart: file, recorded_at, asked_by, prompt_topic).
GET /api/memos/{id}One memo with transcript segments.
PATCH /api/memos/{id}Rate, retitle or correct the year.
GET /api/memos/{id}/audioThe recording, with byte range support.
POST /api/askAsk a question. Streams server sent events: plan, tool, sources, token, removed_quotes, done.
POST /api/ttsRead text aloud (MP3).
GET /api/plannerThe TabPFN forecast.
POST /api/sessions/importImport a call log CSV.
GET /bookThe printable recipe book.

Every call can carry an x-space header: main for the archive, try for the open sandbox. In main, anything that changes data needs the x-family-key header when a passcode is set. GET /api/access tells you whether you can write.

Privacy

On a laptop with Ollama and Whisper, transcription, filing and answering happen on the machine. Only the database is remote, and Atlas can be swapped for a local MongoDB if you want everything offline (vector search then needs Atlas or a local Atlas deployment).

The hosted version sends audio to ElevenLabs for transcription and transcripts to Google AI Studio for Gemma. That trade is made for a demo anyone can open. For a real family archive, run it locally.

Voice cloning is never turned on by default. Only do it with the person's clear permission.

Known limits

  • Uploads on Vercel are capped at 4 MB per file, which is still over 30 minutes of a WhatsApp voice note.
  • Whisper handles English, Yoruba and Hausa unevenly and does not cover Igbo. Use ElevenLabs Scribe for languages it struggles with.
  • The local 4B model is slower and less careful than the hosted 26B. The quote check catches invented quotes from either.
  • The planner is only as good as the ratings. With fewer than 8 it does not guess.