open-dictate Hold a key, speak, let go. The words land at your cursor.
Local-first dictation for Traditional Chinese on Apple Silicon Macs. MLX Whisper turns your voice into text on your own machine, a glossary you own corrects the words it knows, and the sentence is typed where your cursor is. The system may flag and suggest. It may not silently rewrite what you meant.
Early public seed · daemon 0.6.1 · MIT · macOS 14+ · Apple Silicon
01 · Demo One dictation, step by step
週會筆記
幫我把 TouchDesigner 的檔案放上 GitHub 的專案頁,明天下午三點前寄給大家。
“Put the TouchDesigner files on the GitHub project page and send them to everyone before 3 p.m. tomorrow.” Two pairs you taught fired. The time stays as spoken.
Hold fn to dictate the next line.
Hold the fn key above with mouse, finger or Space, then let go. The sentences are scripted; nothing here listens.
This is a demonstration animation. The sentences are scripted and this page never touches your microphone. The real thing is a menu-bar app and a daemon on your Mac. Read the contract →
02 · How it works
Five steps, one rule
A Swift menu-bar app records while you hold the key and hands a WAV file to a Python daemon that keeps the model warm. The daemon transcribes, corrects and punctuates, then answers with text, and the app types it where your cursor is. After transcription, nothing changes your words except pairs you put in the glossary.
-
01 Record
while you hold the key
Hold fn or right Option. The app records 16 kHz mono audio. Presses under half a second are dropped, and a speech-shape gate catches accidental recordings of room noise.
rec = 3.4 svalues from the demo above -
02 Transcribe
what was said
MLX Whisper runs on Apple Silicon inside a daemon that stays loaded. Guards drop known subtitle-credit hallucinations and cut runaway repetition before anything goes further.
model = large-v3-turbovalues from the demo above -
03 Correct
what you meant
Your glossary is a table of wrong → right pairs. Only pairs that match are applied; the sentence is never rewritten. A missed correction is better than a wrong one.
changes = 2values from the demo above -
04 Punctuate
how it reads
The default
smart_zhrules make marks full-width in Chinese context and leave English, numbers and URLs alone. An optional local LLM pass adds fuller punctuation and is thrown away if it changes anything else.punct = "smart_zh"values from the demo above -
05 Insert
where it lands
The text goes in through the Accessibility API, with paste as the fallback. The app never takes focus from the field you were typing in.
insert → AX · pastevalues from the demo above
Beyond the hotkey
- Meeting Mode
- A local recording, or pre-transcribed JSON/JSONL segments, becomes a reviewable package: Markdown, JSONL, SRT and VTT. Same Whisper, same glossary.
- Review queue
- A scanner flags likely mishearings and numbers as candidates. You accept, edit or reject; only accepted pairs reach the glossary, and every decision can be undone.
- Speaker labels
- Transcripts use anonymous labels,
SPEAKER_00,SPEAKER_01. Knowing who is speaking is planned as an optional, local-only layer. - Menu bar
- Teach a pair from the last sentence or from selected text, report a mishearing, and choose the hotkey, the microphone and the punctuation mode.
Seven rules
- Replace words, never sentences. Glossary pairs may correct a word. The system may not rewrite the sentence.
- A missed correction beats a wrong one. One bad pair would poison every transcript after it, so risky pairs wait for context or review.
- Normalise the script, keep the meaning. Traditional Chinese normalisation is allowed. Numbers are not changed automatically.
- Audio stays local. Recording, transcription and correction all run on your Mac.
- Never steal focus. The interface does not pull focus from the text field you are working in.
- Voices are biometric data. Speaker profiles and voice embeddings stay local and are never committed.
- Grow by review. The glossary improves from suggestions you accept, never by silent self-mutation.
Words used on this page
- Push-to-talk
- Hold a key while you speak, release to finish. Open Dictate uses fn or right Option.
- Daemon
dictated.py, the background process that keeps the model loaded so a short sentence needs no cold start. The app talks to it over a Unix socket.- MLX Whisper
- The Whisper speech model running on Apple’s MLX framework. Default:
whisper-large-v3-turbo. - Glossary pair
- One
wrong → rightentry, such asgit hub → GitHub. Applied exactly as written wherever the wrong side appears. - Contextual pair
- A pair that is only safe in some sentences. It is never applied blindly; it waits for context or for you.
- Review queue
- Candidate pairs waiting for accept, edit or reject. Decisions are kept in a history and can be undone.
- smart_zh
- The default, rule-based punctuation: half-width marks after Chinese become full-width. Deterministic and idempotent.
- No-rewrite gate
- The check an optional LLM output must pass: with punctuation removed, it must match the input except for authorised glossary pairs.
- SPEAKER_00
- An anonymous speaker label in a meeting transcript. When the system is unsure who spoke, it says
unknownrather than guess a name.
03 · Status
What works today
Open Dictate is an early public seed, extracted from a private tool used every day. The dictation path works on Apple Silicon Macs; meeting, review and speaker layers are conservative on purpose. Every status below is copied from the README.
Features
- Push-to-talk dictationPublic seed
- Local MLX Whisper daemonPublic seed
- Deterministic glossary correctionPublic seed
- Traditional Chinese full-width punctuationPublic seed
- Menu-bar teaching flowPublic seed
- Meeting transcript packagePublic seedaudio, or JSON/JSONL segments
- Local audio transcription for Meeting ModePublic seedMLX Whisper adapter
- Post-transcription mishearing checkMVPreview flags
- Review-first glossary growthMVPcommand-line queue
- Anonymous speaker labelsMVP
- Local speaker identityPlannedsensitive, optional layer
Know this before you install
- Apple Silicon and macOS 14 or later only. Intel Macs are not supported.
- It installs from source. You need Git, Xcode Command Line Tools and Python 3.11+. There is no DMG yet.
- The installed app depends on your clone. If you move the folder, run the doctor and install again.
- The app is ad-hoc signed. After a rebuild, macOS may ask you to switch its permissions off and on again.
- The first run downloads the model and warming it up can take about two minutes.
On the way to an ordinary Mac app
- Phase 0Release architecture. One product config, data in Application Support, nothing at runtime that needs the clone.
- Phase 1Self-contained beta. Python inside the app, a welcome flow for permissions and the model, a test DMG. Goal: first dictation in ten minutes on a clean Mac, Finder only.
- Phase 2Signed and notarized. Developer ID, Apple notarization, a DMG that opens without Gatekeeper warnings, with checksums.
- Phase 3Updates and rollback. Signed in-app updates, Stable and Beta channels, automatic rollback that never touches your data.
- Phase 4More channels. Homebrew Cask; the Mac App Store only after the sandbox questions are answered.
04 · Privacy
Your voice stays on your Mac
Audio, transcripts, logs, review queues and speaker profiles stay on your machine unless you export them yourself. Dictation needs no server and no account.
Where your data lives
| Data | Location | In the public repo? |
|---|---|---|
| Dictation logs | ~/.open-dictate/dictation-log/ | No |
| Meeting transcripts | ~/.open-dictate/meetings/ | No |
| Review queue | ~/.open-dictate/review-queue/ | No |
| Personal glossary | ~/.open-dictate/glossaries/ | Not by default |
| Speaker profiles | ~/.open-dictate/speakers/ | Never |
| Public test data | fixtures/ · examples/ | Yes, fictional only |
When it does use the network
- Installing. Python packages come from PyPI.
- The model. The first run downloads
mlx-community/whisper-large-v3-turbofrom Hugging Face. When the daemon loads it, the download library may check for a newer copy; that request names the model and carries none of your audio or text. - LLM punctuation, if you turn it on. It talks to Ollama at
127.0.0.1, a server on your own Mac. If you pointOPEN_DICTATE_PUNCT_LLM_URLat another machine, your text goes there. - Meeting Mode with
--model. A model you name is downloaded the first time.
That is the whole list. Once the model is on disk, dictation keeps working without a connection.
Speaker embeddings and voiceprints are biometric data.
05 · Install
Install it
Today Open Dictate installs from source with one script. You need an Apple Silicon Mac on macOS 14 or later, Xcode Command Line Tools and Python 3.11 or newer.
git clone https://github.com/frank890417/open-dictate.git
cd open-dictate
./install.sh # venv, app, LaunchAgents, model warm-up
The installer creates a virtual environment, builds OpenDictate.app into Applications, writes two LaunchAgents for your clone and warms the model. It reports success only after a real ping to the daemon; a socket file alone does not count.
The first run downloads the model, which can take about two minutes. On a slow connection: OPEN_DICTATE_WARM_TIMEOUT=300 ./install.sh.
Then three permissions and one key
- System Settings → Privacy & Security: allow OpenDictate under Microphone, Accessibility and Input Monitoring.
- System Settings → Keyboard: set “Press fn key to” to Do Nothing. Turn off Apple Dictation’s fn shortcut and quit other dictation apps that use the key.
- Click into any text field, hold fn, speak, let go. Prefer right Option? Switch the hotkey from the menu bar.
When something is off
./scripts/doctor.sh # read-only checks, then a real ping
./uninstall.sh # removes the app and LaunchAgents, keeps ~/.open-dictate
./uninstall.sh --purge-data # shows where your data is; never deletes it
The doctor checks the app, its signature, the LaunchAgents, the clone paths, the socket and a real ping. It changes nothing. It can only remind you about permissions, because it does not read the protected macOS permission database.
Meetings and the glossary, from the terminal
# Meeting Mode: a local recording in, a reviewable package out
python3 daemon/meeting_cli.py transcribe ~/Desktop/meeting.m4a --out /tmp/od-meeting --language zh
# the review queue: look, then decide
python3 daemon/glossary/cli.py candidates
python3 daemon/glossary/cli.py accept cand_xxxxx
Mixed English and Chinese? Use --language auto. Output: Markdown, JSONL, SRT, VTT and a JSON summary.
06 · Develop
Build on it
Two processes and one contract. The Swift app owns the hotkey, the microphone and text insertion; the Python daemon owns transcription and correction. They talk in newline-delimited JSON over a Unix socket, and IO-CONTRACT.md is the source of truth for both.
Two lanes
hold fn
|
OpenDictate.app (Swift) records 16 kHz mono PCM16 WAV
| {"cmd": "transcribe", "wav": "/tmp/...wav", "punct": "smart_zh"}
v
/tmp/open-dictate.sock newline-delimited JSON
|
dictated.py (Python, warm)
|-- MLX Whisper audio -> raw text
|-- glossary pairs raw -> corrected (muse_lexicon)
|-- punctuation smart_zh | llm_zh (gated) | raw
|-- log ~/.open-dictate/dictation-log/
v
{"ok": true, "text": "...", "raw": "...", "changes": [...]}
|
OpenDictate.app -> Accessibility insert (paste fallback) -> cursor
meeting.m4a or segments.json / .jsonl
|
meeting_cli.py transcribe
|-- MLX Whisper audio files only
|-- glossary pairs the same deterministic table
|-- speaker labels SPEAKER_00, SPEAKER_01 (anonymous)
|-- QA flags possible mishearings, numbers
v
transcript.md .jsonl .srt .vtt meeting-result.json
The socket
Five commands. Readers ignore fields they do not know, so the daemon can add diagnostics without breaking the app. no_speech is a normal outcome, not an error.
# requests: one JSON object per line on /tmp/open-dictate.sock
{"cmd": "transcribe", "wav": "/tmp/open-dictate-rec-....wav", "punct": "smart_zh"}
{"cmd": "ping"}
{"cmd": "reload_lexicon"}
{"cmd": "add_pair", "wrong": "誤聽", "right": "正確", "source": "dictate-ui"}
{"cmd": "stats"}
# responses
{"ok": true, "text": "校正後文字", "raw": "whisper 原始輸出", "changes": [["誤聽", "正確"]], "punct": "smart_zh"}
{"ok": false, "error": "no_speech"}
One constant pair is a contract on its own: the app’s transcription timeout and the daemon’s punctuation budget. Raise one without the other and the app reports the daemon offline while it finishes in the background. The contract explains the formula.
Where things are
| Part | Path |
|---|---|
| Swift menu-bar app | OpenDictate/ |
| Dictation daemon | daemon/dictated.py |
| Meeting CLI | daemon/meeting_cli.py |
| Mishearing scanner | daemon/qa/mishear_detector.py |
| Review queue | daemon/glossary/ |
| Speaker layer | daemon/speaker/ |
| Starter glossaries | vendor/tools/td-subtitle/glossaries/ |
| Glossary engine | vendor/tools/muse-lexicon/muse_lexicon.py |
| Dictation log (local) | ~/.open-dictate/dictation-log/ |
Build and test
./build.sh
python3 -m unittest discover tests
python3 scripts/golden-bench.py --skip-daemon
python3 scripts/public-safety-scan.py
./scripts/smoke-test.sh # builds, tests, exports a meeting demo, checks the signed app
Before a pull request
- Keep examples fictional or public domain.
- Never commit real audio, transcripts, dictation logs, private paths, personal glossaries or speaker profiles.
- Prefer deterministic correction and review queues to silent rewriting.
- Run the safety scan and the tests; run
./build.shif you touched Swift.
Docs
Documents mix English and Traditional Chinese; a few are in one language only.
- IO-CONTRACT.mdPaths, the socket protocol, glossary schema, correction rules, quality gates
- SetupRequirements, permissions, hotkey conflicts, doctor, uninstall
- Privacy modelData classes, where they live, what may never be committed
- Meeting ModeInputs, the pipeline, the segment schema, the safety rule
- Self-evolving glossaryThe review loop, its commands and its four buckets
- Speaker identityAnonymous labels now, optional local profiles later
- macOS distributionFrom install script to a signed, notarized, updatable app
- MCP roadmapA local, consent-gated MCP server, phase by phase
- ContributingGround rules and the commands to run
07 · For agents
For AI agents
A coding agent can install, diagnose and extend Open Dictate from a terminal, and everything it needs is plain text: this page, /llms.txt and the contract. What it cannot do is grant macOS permissions; those stay with you.
Install on this Mac: paste into Claude Code
Clone https://github.com/frank890417/open-dictate and read README.md, docs/SETUP.md and docs/PRIVACY.md first.
Check that this Mac is Apple Silicon on macOS 14 or later (uname -m, sw_vers). If not, stop and tell me.
Run ./install.sh. Then run ./scripts/doctor.sh and python3 daemon/dictate_cli.py ping, and report what they say.
Tell me which three permissions to allow in System Settings and how to set the fn key. You cannot grant them for me.
Do not delete anything under ~/.open-dictate. It holds my glossary and logs.
Change the code: paste into Claude Code
Clone https://github.com/frank890417/open-dictate. Read IO-CONTRACT.md before touching daemon/ or OpenDictate/; it is the contract between them.
Task: <describe the change, e.g. "add a field to the stats response">
Rules from the repo: correction is glossary pairs only, never sentence rewriting; prefer a missed correction to a wrong one; keep the wire protocol backward compatible; use fictional data in fixtures and tests.
Done when: python3 -m unittest discover tests, python3 scripts/golden-bench.py --skip-daemon and python3 scripts/public-safety-scan.py pass, and ./build.sh succeeds if Swift changed.
Machine-readable
- /llms.txtWhat this is, the pipeline, the rules, every doc, in llmstxt.org format
- /llms-full.txtREADME, IO-CONTRACT.md and every doc in one file
- IO-CONTRACT.mdThe source of truth between the app, the daemon and the glossary engine
application/ld+jsonThis page carries SoftwareApplication JSON-LD- /zh/This page in Traditional Chinese
MCP: planned, not shipped
There is no MCP server yet. The roadmap starts with a local, read-only stdio server for status, capabilities and glossary counts, and adds actions only behind consent given at the moment of use. Recording from the microphone stays an app interaction: no background or blanket consent.
08 · Lineage
Where it came from
Open Dictate was extracted from a private dictation tool used every day, then rebuilt as a public-safe repository with no history: no recordings, no transcripts, no personal glossary, only the engine, the contract and fictional examples.
The public repository is now the single source of truth. The author’s private build keeps only personal settings and adapters on top of this same core and packaging pipeline. Its sibling, open-audiovisual, was opened the same way: a tool that served real work first.
The system may flag and suggest, but it must not silently rewrite meaning.
Next, from the roadmap
- Local speaker profiles, enrolled with consent, stored only on your Mac
- A menu-bar view for reviewing glossary candidates
- Import and export for glossary packages you own
- An app that runs without the clone, developer Python or Xcode tools
- A Developer ID signed, notarized DMG tested on clean Macs
- Signed in-app updates, rollback, Stable and Beta channels