open-dictate
GitHub

open-dictate Hold a key, speak, let go. The words land at your cursor.

Local-first dictation for Traditional Chinese on Apple Silicon Macs. MLX Whisper turns your voice into text on your own machine, a glossary you own corrects the words it knows, and the sentence is typed where your cursor is. The system may flag and suggest. It may not silently rewrite what you meant.

Early public seed · daemon 0.6.1 · MIT · macOS 14+ · Apple Silicon

01 · Demo One dictation, step by step

Any text field, any appdemonstration

週會筆記

幫我把 TouchDesigner 的檔案放上 GitHub 的專案頁,明天下午三點前寄給大家。

“Put the TouchDesigner files on the GitHub project page and send them to everyone before 3 p.m. tomorrow.” Two pairs you taught fired. The time stays as spoken.

3.4 s

Hold fn to dictate the next line.

Hold the fn key above with mouse, finger or Space, then let go. The sentences are scripted; nothing here listens.

This is a demonstration animation. The sentences are scripted and this page never touches your microphone. The real thing is a menu-bar app and a daemon on your Mac. Read the contract →

02 · How it works

Five steps, one rule

A Swift menu-bar app records while you hold the key and hands a WAV file to a Python daemon that keeps the model warm. The daemon transcribes, corrects and punctuates, then answers with text, and the app types it where your cursor is. After transcription, nothing changes your words except pairs you put in the glossary.

  1. 01 Record

    while you hold the key

    Hold fn or right Option. The app records 16 kHz mono audio. Presses under half a second are dropped, and a speech-shape gate catches accidental recordings of room noise.

    rec = 3.4 svalues from the demo above

  2. 02 Transcribe

    what was said

    MLX Whisper runs on Apple Silicon inside a daemon that stays loaded. Guards drop known subtitle-credit hallucinations and cut runaway repetition before anything goes further.

    model = large-v3-turbovalues from the demo above

  3. 03 Correct

    what you meant

    Your glossary is a table of wrong → right pairs. Only pairs that match are applied; the sentence is never rewritten. A missed correction is better than a wrong one.

    changes = 2values from the demo above

  4. 04 Punctuate

    how it reads

    The default smart_zh rules make marks full-width in Chinese context and leave English, numbers and URLs alone. An optional local LLM pass adds fuller punctuation and is thrown away if it changes anything else.

    punct = "smart_zh"values from the demo above

  5. 05 Insert

    where it lands

    The text goes in through the Accessibility API, with paste as the fallback. The app never takes focus from the field you were typing in.

    insert → AX · pastevalues from the demo above

Beyond the hotkey

Meeting Mode
A local recording, or pre-transcribed JSON/JSONL segments, becomes a reviewable package: Markdown, JSONL, SRT and VTT. Same Whisper, same glossary.
Review queue
A scanner flags likely mishearings and numbers as candidates. You accept, edit or reject; only accepted pairs reach the glossary, and every decision can be undone.
Speaker labels
Transcripts use anonymous labels, SPEAKER_00, SPEAKER_01. Knowing who is speaking is planned as an optional, local-only layer.
Menu bar
Teach a pair from the last sentence or from selected text, report a mishearing, and choose the hotkey, the microphone and the punctuation mode.

Seven rules

  1. Replace words, never sentences. Glossary pairs may correct a word. The system may not rewrite the sentence.
  2. A missed correction beats a wrong one. One bad pair would poison every transcript after it, so risky pairs wait for context or review.
  3. Normalise the script, keep the meaning. Traditional Chinese normalisation is allowed. Numbers are not changed automatically.
  4. Audio stays local. Recording, transcription and correction all run on your Mac.
  5. Never steal focus. The interface does not pull focus from the text field you are working in.
  6. Voices are biometric data. Speaker profiles and voice embeddings stay local and are never committed.
  7. Grow by review. The glossary improves from suggestions you accept, never by silent self-mutation.

Words used on this page

Push-to-talk
Hold a key while you speak, release to finish. Open Dictate uses fn or right Option.
Daemon
dictated.py, the background process that keeps the model loaded so a short sentence needs no cold start. The app talks to it over a Unix socket.
MLX Whisper
The Whisper speech model running on Apple’s MLX framework. Default: whisper-large-v3-turbo.
Glossary pair
One wrong → right entry, such as git hub → GitHub. Applied exactly as written wherever the wrong side appears.
Contextual pair
A pair that is only safe in some sentences. It is never applied blindly; it waits for context or for you.
Review queue
Candidate pairs waiting for accept, edit or reject. Decisions are kept in a history and can be undone.
smart_zh
The default, rule-based punctuation: half-width marks after Chinese become full-width. Deterministic and idempotent.
No-rewrite gate
The check an optional LLM output must pass: with punctuation removed, it must match the input except for authorised glossary pairs.
SPEAKER_00
An anonymous speaker label in a meeting transcript. When the system is unsure who spoke, it says unknown rather than guess a name.

03 · Status

What works today

Open Dictate is an early public seed, extracted from a private tool used every day. The dictation path works on Apple Silicon Macs; meeting, review and speaker layers are conservative on purpose. Every status below is copied from the README.

Features

  • Push-to-talk dictationPublic seed
  • Local MLX Whisper daemonPublic seed
  • Deterministic glossary correctionPublic seed
  • Traditional Chinese full-width punctuationPublic seed
  • Menu-bar teaching flowPublic seed
  • Meeting transcript packagePublic seedaudio, or JSON/JSONL segments
  • Local audio transcription for Meeting ModePublic seedMLX Whisper adapter
  • Post-transcription mishearing checkMVPreview flags
  • Review-first glossary growthMVPcommand-line queue
  • Anonymous speaker labelsMVP
  • Local speaker identityPlannedsensitive, optional layer

Know this before you install

  • Apple Silicon and macOS 14 or later only. Intel Macs are not supported.
  • It installs from source. You need Git, Xcode Command Line Tools and Python 3.11+. There is no DMG yet.
  • The installed app depends on your clone. If you move the folder, run the doctor and install again.
  • The app is ad-hoc signed. After a rebuild, macOS may ask you to switch its permissions off and on again.
  • The first run downloads the model and warming it up can take about two minutes.

On the way to an ordinary Mac app

  1. Phase 0Release architecture. One product config, data in Application Support, nothing at runtime that needs the clone.
  2. Phase 1Self-contained beta. Python inside the app, a welcome flow for permissions and the model, a test DMG. Goal: first dictation in ten minutes on a clean Mac, Finder only.
  3. Phase 2Signed and notarized. Developer ID, Apple notarization, a DMG that opens without Gatekeeper warnings, with checksums.
  4. Phase 3Updates and rollback. Signed in-app updates, Stable and Beta channels, automatic rollback that never touches your data.
  5. Phase 4More channels. Homebrew Cask; the Mac App Store only after the sandbox questions are answered.

The distribution roadmap →

04 · Privacy

Your voice stays on your Mac

Audio, transcripts, logs, review queues and speaker profiles stay on your machine unless you export them yourself. Dictation needs no server and no account.

Where your data lives

DataLocationIn the public repo?
Dictation logs~/.open-dictate/dictation-log/No
Meeting transcripts~/.open-dictate/meetings/No
Review queue~/.open-dictate/review-queue/No
Personal glossary~/.open-dictate/glossaries/Not by default
Speaker profiles~/.open-dictate/speakers/Never
Public test datafixtures/ · examples/Yes, fictional only

When it does use the network

  1. Installing. Python packages come from PyPI.
  2. The model. The first run downloads mlx-community/whisper-large-v3-turbo from Hugging Face. When the daemon loads it, the download library may check for a newer copy; that request names the model and carries none of your audio or text.
  3. LLM punctuation, if you turn it on. It talks to Ollama at 127.0.0.1, a server on your own Mac. If you point OPEN_DICTATE_PUNCT_LLM_URL at another machine, your text goes there.
  4. Meeting Mode with --model. A model you name is downloaded the first time.

That is the whole list. Once the model is on disk, dictation keeps working without a connection.

Speaker embeddings and voiceprints are biometric data.

Keep them local, never commit them, never paste them into issues, and enroll a voice only with consent. Share transcripts with anonymous labels.

The full privacy model →

05 · Install

Install it

Today Open Dictate installs from source with one script. You need an Apple Silicon Mac on macOS 14 or later, Xcode Command Line Tools and Python 3.11 or newer.

terminal
git clone https://github.com/frank890417/open-dictate.git
cd open-dictate
./install.sh        # venv, app, LaunchAgents, model warm-up

The installer creates a virtual environment, builds OpenDictate.app into Applications, writes two LaunchAgents for your clone and warms the model. It reports success only after a real ping to the daemon; a socket file alone does not count.

The first run downloads the model, which can take about two minutes. On a slow connection: OPEN_DICTATE_WARM_TIMEOUT=300 ./install.sh.

Then three permissions and one key

  1. System Settings → Privacy & Security: allow OpenDictate under Microphone, Accessibility and Input Monitoring.
  2. System Settings → Keyboard: set “Press fn key to” to Do Nothing. Turn off Apple Dictation’s fn shortcut and quit other dictation apps that use the key.
  3. Click into any text field, hold fn, speak, let go. Prefer right Option? Switch the hotkey from the menu bar.

When something is off

terminal
./scripts/doctor.sh            # read-only checks, then a real ping
./uninstall.sh                 # removes the app and LaunchAgents, keeps ~/.open-dictate
./uninstall.sh --purge-data    # shows where your data is; never deletes it

The doctor checks the app, its signature, the LaunchAgents, the clone paths, the socket and a real ping. It changes nothing. It can only remind you about permissions, because it does not read the protected macOS permission database.

Meetings and the glossary, from the terminal

terminal
# Meeting Mode: a local recording in, a reviewable package out
python3 daemon/meeting_cli.py transcribe ~/Desktop/meeting.m4a --out /tmp/od-meeting --language zh

# the review queue: look, then decide
python3 daemon/glossary/cli.py candidates
python3 daemon/glossary/cli.py accept cand_xxxxx

Mixed English and Chinese? Use --language auto. Output: Markdown, JSONL, SRT, VTT and a JSON summary.

Full setup guide →

06 · Develop

Build on it

Two processes and one contract. The Swift app owns the hotkey, the microphone and text insertion; the Python daemon owns transcription and correction. They talk in newline-delimited JSON over a Unix socket, and IO-CONTRACT.md is the source of truth for both.

Two lanes

dictation
hold fn
  |
OpenDictate.app (Swift)       records 16 kHz mono PCM16 WAV
  |  {"cmd": "transcribe", "wav": "/tmp/...wav", "punct": "smart_zh"}
  v
/tmp/open-dictate.sock        newline-delimited JSON
  |
dictated.py (Python, warm)
  |-- MLX Whisper             audio -> raw text
  |-- glossary pairs          raw -> corrected   (muse_lexicon)
  |-- punctuation             smart_zh | llm_zh (gated) | raw
  |-- log                     ~/.open-dictate/dictation-log/
  v
{"ok": true, "text": "...", "raw": "...", "changes": [...]}
  |
OpenDictate.app  ->  Accessibility insert (paste fallback)  ->  cursor
meeting mode
meeting.m4a   or   segments.json / .jsonl
  |
meeting_cli.py transcribe
  |-- MLX Whisper             audio files only
  |-- glossary pairs          the same deterministic table
  |-- speaker labels          SPEAKER_00, SPEAKER_01 (anonymous)
  |-- QA flags                possible mishearings, numbers
  v
transcript.md  .jsonl  .srt  .vtt  meeting-result.json

The socket

Five commands. Readers ignore fields they do not know, so the daemon can add diagnostics without breaking the app. no_speech is a normal outcome, not an error.

IO-CONTRACT.md
# requests: one JSON object per line on /tmp/open-dictate.sock
{"cmd": "transcribe", "wav": "/tmp/open-dictate-rec-....wav", "punct": "smart_zh"}
{"cmd": "ping"}
{"cmd": "reload_lexicon"}
{"cmd": "add_pair", "wrong": "誤聽", "right": "正確", "source": "dictate-ui"}
{"cmd": "stats"}

# responses
{"ok": true, "text": "校正後文字", "raw": "whisper 原始輸出", "changes": [["誤聽", "正確"]], "punct": "smart_zh"}
{"ok": false, "error": "no_speech"}

One constant pair is a contract on its own: the app’s transcription timeout and the daemon’s punctuation budget. Raise one without the other and the app reports the daemon offline while it finishes in the background. The contract explains the formula.

Where things are

PartPath
Swift menu-bar appOpenDictate/
Dictation daemondaemon/dictated.py
Meeting CLIdaemon/meeting_cli.py
Mishearing scannerdaemon/qa/mishear_detector.py
Review queuedaemon/glossary/
Speaker layerdaemon/speaker/
Starter glossariesvendor/tools/td-subtitle/glossaries/
Glossary enginevendor/tools/muse-lexicon/muse_lexicon.py
Dictation log (local)~/.open-dictate/dictation-log/

Build and test

terminal
./build.sh
python3 -m unittest discover tests
python3 scripts/golden-bench.py --skip-daemon
python3 scripts/public-safety-scan.py
./scripts/smoke-test.sh        # builds, tests, exports a meeting demo, checks the signed app

Before a pull request

  • Keep examples fictional or public domain.
  • Never commit real audio, transcripts, dictation logs, private paths, personal glossaries or speaker profiles.
  • Prefer deterministic correction and review queues to silent rewriting.
  • Run the safety scan and the tests; run ./build.sh if you touched Swift.

Docs

Documents mix English and Traditional Chinese; a few are in one language only.

07 · For agents

For AI agents

A coding agent can install, diagnose and extend Open Dictate from a terminal, and everything it needs is plain text: this page, /llms.txt and the contract. What it cannot do is grant macOS permissions; those stay with you.

Install on this Mac: paste into Claude Code

prompt
Clone https://github.com/frank890417/open-dictate and read README.md, docs/SETUP.md and docs/PRIVACY.md first.
Check that this Mac is Apple Silicon on macOS 14 or later (uname -m, sw_vers). If not, stop and tell me.
Run ./install.sh. Then run ./scripts/doctor.sh and python3 daemon/dictate_cli.py ping, and report what they say.
Tell me which three permissions to allow in System Settings and how to set the fn key. You cannot grant them for me.

Do not delete anything under ~/.open-dictate. It holds my glossary and logs.

Change the code: paste into Claude Code

prompt
Clone https://github.com/frank890417/open-dictate. Read IO-CONTRACT.md before touching daemon/ or OpenDictate/; it is the contract between them.

Task: <describe the change, e.g. "add a field to the stats response">

Rules from the repo: correction is glossary pairs only, never sentence rewriting; prefer a missed correction to a wrong one; keep the wire protocol backward compatible; use fictional data in fixtures and tests.
Done when: python3 -m unittest discover tests, python3 scripts/golden-bench.py --skip-daemon and python3 scripts/public-safety-scan.py pass, and ./build.sh succeeds if Swift changed.

Machine-readable

  • /llms.txtWhat this is, the pipeline, the rules, every doc, in llmstxt.org format
  • /llms-full.txtREADME, IO-CONTRACT.md and every doc in one file
  • IO-CONTRACT.mdThe source of truth between the app, the daemon and the glossary engine
  • application/ld+jsonThis page carries SoftwareApplication JSON-LD
  • /zh/This page in Traditional Chinese

MCP: planned, not shipped

There is no MCP server yet. The roadmap starts with a local, read-only stdio server for status, capabilities and glossary counts, and adds actions only behind consent given at the moment of use. Recording from the microphone stays an app interaction: no background or blanket consent.

The MCP roadmap →

08 · Lineage

Where it came from

Open Dictate was extracted from a private dictation tool used every day, then rebuilt as a public-safe repository with no history: no recordings, no transcripts, no personal glossary, only the engine, the contract and fictional examples.

The public repository is now the single source of truth. The author’s private build keeps only personal settings and adapters on top of this same core and packaging pipeline. Its sibling, open-audiovisual, was opened the same way: a tool that served real work first.

The system may flag and suggest, but it must not silently rewrite meaning.

Design rule · README

Next, from the roadmap

  1. Local speaker profiles, enrolled with consent, stored only on your Mac
  2. A menu-bar view for reviewing glossary candidates
  3. Import and export for glossary packages you own
  4. An app that runs without the clone, developer Python or Xcode tools
  5. A Developer ID signed, notarized DMG tested on clean Macs
  6. Signed in-app updates, rollback, Stable and Beta channels

Roadmap in the README →