Skip to main content

Settings → Voice

Everything Mnemosyne OS does with sound is set here: how it hears you (microphone and transcription), how it speaks (the read-aloud voice), and how the ∞ voice assistant behaves. The rail entry is Voice; the page subtitle reads Voice input (dictation) and voice output (read-aloud, incl. ElevenLabs).

Every control on this page saves as soon as you touch it. There is no Save button.

For what the voice does once it is set up, see Calliope, the voice. This page is about the settings.

Microphone

The first cause of a garbage transcription (a reply of just « you ») is a silent microphone, so the page opens with a live meter.

  • A device picker lists System default microphone and every input the OS reports. Device names appear once the microphone permission has been granted.
  • Test microphone starts the meter; the button becomes Stop test. While it runs, the page says Speak now — the bar should move. The bar turns grey below a whisper, amber at a low level, green when the signal is good, and reports the Peak reached.
  • No signal detected — wrong input device or muted mic? means the peak never left zero: switch device in the picker, or unmute.
  • Microphone access denied. Allow it in Windows → Privacy → Microphone. is the OS refusing, not the app.

Nothing is recorded during the test. It is a level meter and nothing else.

Default engine (transcription)

Two cards choose how your speech becomes text:

  • Local (Whisper): Offline, private. Needs a downloaded model. Below the card, a line names the Active model and whether it is ready to use or not downloaded — open Advanced to get it.
  • Cloud: Ready · openai (or gemini) when your active AI provider can transcribe, otherwise Needs an OpenAI/Gemini key. With the card selected, a line says Transcription via your configured model: OpenAI whisper-1 (or Gemini). If your active cloud model is one that does not transcribe audio, the line turns amber: Your active cloud model can't transcribe audio. Set OpenAI or Gemini in AI & Models.
Local wins when cloud is not set up

If you pick Cloud but no key can transcribe, a downloaded local model is used anyway. The choice here is a preference, never a way to lose dictation.

Advanced — manage models

Collapsed by default; it opens itself while a download runs so the progress stays visible. Two tiers:

Local models (Whisper), which run inside the app:

ModelSizeThe app's own note
Whisper Tiny~75 MBFastest, smallest — lower accuracy. Multilingual.
Whisper Base~150 MBBalanced speed/accuracy. Multilingual.
Whisper Small~250 MBGreat accuracy, reasonable speed. Multilingual.
Distil-Whisper (EN)~180 MBFast & accurate, English only.

Large models (GPU speech engine), which run in a separate process:

ModelSizeThe app's own note
Whisper Medium~1.5 GBHigh accuracy, multilingual.
Whisper Large v3 Turbo ★~1.6 GBNear large-v3 accuracy, much faster. Recommended.
Whisper Large v3~3 GBMaximum accuracy, multilingual. The heaviest model.

Click a row to make it the active model; Download fetches it. For a large model, the first Download installs the GPU speech engine first (the button reads Installing 40 %, then Downloading 40 %). The status line above the tier explains the machine you are on:

  • Downloading a model below sets up the GPU engine automatically the first time (one-time, needs Python 3.10+). GPU acceleration requires an NVIDIA graphics card (CUDA) — other machines transcribe on the CPU.
  • GPU speech engine installed · CUDA once it runs on the card.
  • NVIDIA GPU detected — enable acceleration for much faster transcription. with an Enable GPU button, when the card is there but the CUDA libraries are not yet (about 800 MB, downloaded once).
  • no NVIDIA GPU: runs on CPU (slower). on every other machine. The large models still work there, only slower.

A downloaded model shows Downloaded and a bin icon. The bin asks Delete? and offers Delete / Cancel before removing anything.

Voice output (read-aloud)

Three cards choose the voice that reads replies aloud, in the chat, in the ∞ conversation and in MnemoReader:

  • System (offline): Private, free, robotic. The voices built into your OS, nothing to download.
  • Local neural: Natural, free, offline. Downloadable. Engines that run on your machine, described below.
  • ElevenLabs (cloud): Natural voice. Needs a key; text leaves your machine.

ElevenLabs: entering your API key

  1. Create a key on ElevenLabs

    In your ElevenLabs account, open the API keys page (it lives under your profile menu) and create a key. Copy it right away: ElevenLabs shows it once. If you restrict what the key may do, it needs text-to-speech and read access to your voices.

  2. Select the ElevenLabs (cloud) card

    Open Settings → Voice, scroll to Voice output (read-aloud) and click ElevenLabs (cloud). A field labelled ElevenLabs API key appears under the cards.

  3. Paste the key

    Paste it into the field. It is a password field, so the characters stay hidden, and it saves as you type. There is no Save button to press.

  4. Load voices

    Click Load voices. The button is greyed until the field holds something. The voices of your account appear in a dropdown next to it, with the first one already selected. If the key is wrong, the error ElevenLabs returned is printed in red under the field.

  5. Test voice

    Once a voice is selected, Test voice speaks a full sentence in the language of your interface, at the read-aloud volume set lower on the page.

  6. Choose the model

    Two buttons pick the ElevenLabs model: Quality (Richest voice · 1 credit/char) or Eco + fast (~½ the cost · lower latency). The credits are ElevenLabs credits, billed by ElevenLabs.

The key stays on your machine, in the app's own settings, and is only ever sent to ElevenLabs, with each request. The card says it plainly: with this engine, the text of every reply read aloud leaves your machine.

Local neural

With Local neural selected, a line states which engine is speaking (Reading with: Piper), then one card per engine:

EngineSizeThe app's own note
Piper~290 MBNatural French, Spanish & English · offline
XTTS-v2 (voice cloning)~5 GBClones your voice · FR/ES/EN · GPU · needs Python
Chatterbox~4,5 GBExpressive · 23 languages · clones your voice · GPU · needs Python
Zonos-v0.1~5,5 GBEmotion vector · FR/ES/EN/DE/ZH/JA · 6 GB VRAM · needs Python

Zonos is greyed and marked Soon: it is parked, not installable.

Engramm license

Local neural voices are part of the Engramm license. Without it, every card wears an amber Engramm license badge, and Download stays clickable so that the click can explain why (Local neural voices need an Engramm license.) instead of doing nothing.

On each card:

  • Download installs the engine. The button counts Downloading 42 % · torch with the current stage, because a multi-gigabyte install whose percentage sits still reads as frozen. A red × stops it; a stopped install is not an error. While the app is still asking whether an engine is on disk, the card says Checking… rather than guessing.
  • Use makes an installed engine the one that reads aloud. Clicking a card that is not installed says Download Piper first, then select it.
  • Load / Unload (every engine but Piper) put the model in memory now, so the first sentence does not pay the cold start (Load it now so the first sentence is instant.), or drop it to free the graphics memory (Free its VRAM. The next sentence reloads it.).
  • A voice dropdown and Test voice. For Piper, + More voices (French, English, 40+ languages) opens the Voice library: search, a language filter, a quality filter, a Use button per voice. The ↻ voices button fetches any new voice files.
  • Speed: Piper runs from slower to faster (0.5× to 2×). Chatterbox only slows down (0.5× to 1×, the right end says natural): its engine ignores anything faster, so the slider does not offer it. XTTS has no speed control at all.
  • Expressivity (Chatterbox): flatter to more animated. The note under it matters: Baked into the cloned voice: changing it re-prepares the clone once, so the next sentence takes longer.
  • On a cloning engine, Clone my voice records from the microphone until you tap again (60 s cap, the button counts 23s / 60s · tap to stop) and From audio file uses a recording you already have. The line below says what is on disk, with its length and date (Voice sample saved · 23s · …), or No voice sample recorded yet. One recording serves every cloning engine, and 15–60 s of clean voice clones best. Once you have a sample, the status line warns you when the selected engine cannot use it: this engine does not use your cloned voice. Pick one that does.

When two engines are installed, a box under the cards splits the surfaces:

  • Engine for the ∞ conversation: A conversation cannot wait for a slow engine; read-aloud can. Default Same as read-aloud.
  • Engine for long reading (MnemoReader): A book can wait for a slower engine — nothing is answering you. Pick the one that sounds best.
Two engines in memory

A split keeps two models loaded at once. On a card that cannot hold both, every sentence takes several times longer, and CUDA does not report it. The page says so before the selectors, and shows a red box (Two voice engines are loaded at once…) whenever it actually happens. Unload the one you are not using.

Read-aloud volume

A slider from 0 to 100 %, in steps of 5. Volume when Mnemosyne OS reads replies aloud (speaker button on messages). The Test voice buttons above play at this value, so what you hear is what the chat will use.

Three switches for the ∞ conversation:

  • Talk over to interrupt (barge-in): Headphones required. Speak while it replies and it stops to listen. On speakers it hears itself and cuts off.
  • Expressive eyes: The orb gets blinking eyes that react to its state (listening, thinking, speaking). Off = pure ∞.
  • Generate replies locally (saves cloud tokens): The ∞ assistant writes its answers with its OWN small local model (Qwen 1.5B, downloaded below) — your principal AI (cloud) is untouched. Switching it on reveals a row for that model, Qwen 2.5 1.5B · ~1 GB · the assistant's own brain, with its own Download button.

Orb appearance

A live thumbnail of the current orb and its preset name, and a Customize button that opens the Orb studio:

  • Size in full screen, a slider, and a Full screen switch (The orb fills the whole screen in assistant mode — the size setting no longer applies.).
  • Face presets: Minimal, Soft, Mascot, Kawaii.
  • Abstract orbs: Aurora, Nebula, Plasma, Halo, Constellation. A tile says More via MnemoHub — soon.
  • Fine-tune for the face family: Eyes (Sparkle, Round, Dot), 3D volume, Eyebrows, Blush cheeks. Touch any of them and the preset reads Custom.

Pronunciation and hearing corrections

Two editable tables, each a list of word → replacement rows with Add a word / Add a correction and a × to remove:

  • Pronunciation (word → spoken) rewrites words before any voice speaks them. Fix how specific words are said (names, acronyms). E.g. Mnemosyne → Némozine.
  • Hearing corrections (mis-heard → correct) rewrites every transcript. Fix words the transcriber keeps getting wrong. E.g. Némozine → Mnemosyne.

They apply to every engine, local or cloud.

The line at the bottom

Local models run fully offline. Cloud uses your configured OpenAI/Gemini key (audio leaves your machine). That is the whole privacy model of this page in one sentence: the local tiers never send anything; the cloud transcription sends your audio, and ElevenLabs receives the text it reads.