# Open Dictate: full text for language models > Every document of Open Dictate in one file: README, IO-CONTRACT.md and docs/. > Generated by site/build.mjs from the repository; the short index is /llms.txt. > Site: https://frank890417.github.io/open-dictate/ (English) · https://frank890417.github.io/open-dictate/zh/ (繁體中文) > Repository: https://github.com/frank890417/open-dictate (MIT) --- # Open Dictate 🎙️ > **本地優先的 macOS 繁體中文語音輸入與會議轉錄工具。** > **Local-first Traditional Chinese dictation and meeting transcription for macOS.** **[Website](https://frank890417.github.io/open-dictate/)** · [Setup](docs/SETUP.md) · [Privacy](docs/PRIVACY.md) Hold a hotkey, speak, release, and Open Dictate inserts the transcribed sentence at your cursor. The core path runs locally with MLX Whisper on Apple Silicon, then applies deterministic glossary correction owned by the user. 按住快捷鍵、說話、放開,Open Dictate 會把轉錄後的文字插入目前游標位置。核心流程在 Apple Silicon 上本地執行 MLX Whisper,並套用使用者擁有的確定性詞庫校正。 ![Open Dictate screenshot](docs/assets/open-dictate-screenshot.svg) ## 快速介紹 / TL;DR Open Dictate is not only a Whisper wrapper. The long-term loop is: Open Dictate 不只是 Whisper 包裝器。完整方向是這個閉環: ```text Speak / record → local transcription → deterministic correction → speaker-aware transcript → possible mishearing candidates → human review → personal glossary improves → next transcription gets better ``` Design rule: **the system may flag and suggest, but it must not silently rewrite meaning.** 設計原則:**系統可以標記、可以建議,但不能靜默改寫意思。** ## 功能狀態 / Feature Status | 功能 / Feature | 狀態 / Status | |---|---| | Push-to-talk dictation / 按住說話輸入 | Public seed | | Local MLX Whisper daemon / 本地 Whisper 常駐 daemon | Public seed | | Deterministic glossary correction / 確定性詞庫校正 | Public seed | | Traditional Chinese punctuation / 繁中全形標點 | Public seed | | Menubar teaching flow / 選單列教詞庫 | Public seed | | Meeting transcript package / 會議逐字稿包 | Public seed: audio or JSON/JSONL segments | | Local audio ASR for Meeting Mode / 會議模式本地音檔 ASR | Public seed: MLX Whisper adapter | | Post-transcription QA / 轉錄後誤聽偵測 | MVP: review flags | | Review-first glossary growth / 審核優先詞庫成長 | MVP: CLI queue | | Anonymous speaker labels / 匿名說話者標籤 | MVP | | Local speaker identity / 本機說話者身分 | Planned, sensitive optional layer | ## 隱私模型 / Privacy Model Audio stays local by default. Dictation logs, meeting transcripts, review queues, personal glossaries, and speaker profiles belong to the user and should live under local paths such as `~/.open-dictate/`. 音檔預設留在本機。語音輸入日誌、會議逐字稿、審核佇列、個人詞庫、說話者資料都屬於使用者,應放在 `~/.open-dictate/` 這類本機路徑。 Speaker embeddings and voiceprints are biometric data. Open Dictate's public repo only ships interfaces, schemas, and fictional examples. Real speaker profiles must stay local and must not be committed. 聲紋與說話者 embedding 是生物特徵資料。Open Dictate 的公開 repo 只提供介面、schema 與虛構範例;真實說話者資料必須留在本機,不能 commit。 Read: [`docs/PRIVACY.md`](docs/PRIVACY.md) ## 快速開始 / Quick Start ```bash git clone https://github.com/frank890417/open-dictate.git cd open-dictate ./install.sh # standalone install: creates .venv-dictate and uses vendor starter glossaries ``` The first install downloads Python packages and the Whisper model. Model warm-up can take about two minutes; the installer shows progress and waits up to 180 seconds before printing diagnostics. 第一次安裝會下載 Python 套件與 Whisper 模型。模型載入可能約需兩分鐘;安裝器會顯示進度,最多等待 180 秒,逾時則印出診斷資訊。 First-time permissions / 第一次使用需要開權限: 1. System Settings → Privacy & Security → enable OpenDictate for: - Accessibility - Input Monitoring - Microphone 2. If you use the `fn` hotkey: Keyboard → “Press 🌐 key to” → Do Nothing. Disable Apple Dictation's `fn` shortcut and quit other dictation apps that capture the same key. 3. Click any text field, hold `fn`, speak, release. Diagnostics and uninstall / 診斷與解除安裝: ```bash ./scripts/doctor.sh # read-only checks; does not install or start anything ./uninstall.sh # removes app + LaunchAgents, keeps ~/.open-dictate ./uninstall.sh --purge-data # prints manual data-removal guidance; never deletes user data ``` Full setup: [`docs/SETUP.md`](docs/SETUP.md) ## 即時語音輸入 / Real-time Dictation - Push-to-talk: `fn` or right Option. - 16 kHz mono PCM16 recording → Python daemon → corrected text insertion. - Short-press and silence gates reduce accidental hallucinations. - **Repetition-loop guard (v0.5.3)**: Whisper degenerates into token loops at both ends of the duration range — very short clips and long dictations — repeating one character or phrase until the window ends. A deterministic tail pass removes the runaway and falls back to `no_speech` when the whole utterance is a loop. It is conservative by design: natural emphasis such as `對對對對` is left untouched. - **Length-aware punctuation budget (v0.5.3)**: the optional LLM punctuation pass rewrites the whole segment, so its cost grows with text length. A fixed timeout made long dictations structurally impossible to complete. The budget is now `min(4.0 + chars × 0.025, 10.0)` seconds. Note that length is not the only cause of timeouts — machine-wide contention is at least as common — so the timeout log line now prints characters, budget, elapsed time and machine load. - **Robustness pass (v0.6.0)**: a background warm keeper stops cold-model requests from cancelling the punctuation model load; only contextual pairs present in the sentence are sent to the LLM; the daemon respects the shell deadline; ASR decode guards (`sample_len` by duration, temperature capped at 0.4), a false-trigger gate, a subtitle-credit hallucination filter, a 200-token Whisper prompt with the style sentence at the end, and a fix for a regex that could block the daemon for minutes. Details: IO-CONTRACT §Daemon 0.6.0 robustness notes. - **Punctuation fixes (v0.6.1)**: Whisper's small-form marks (`﹐﹑﹗﹖`) are treated as pauses, not tone (comma mid-sentence, full stop at the end; `﹖` stays `?` only when the clause has a question word), and an enumeration comma right before a coordinating conjunction is removed or turned into a comma after the punctuation layer. Daemon-internal; wire protocol stays 1.0. - Text insertion uses Accessibility direct insertion when possible, then paste fallback. - Optional microphone selection. ## 會議模式 / Meeting Mode Meeting Mode creates a reviewable transcript package from longer conversations. It now accepts local audio files through an MLX Whisper adapter, or pre-transcribed JSON/JSONL segments for integrations and tests. The public-safe default still uses anonymous speaker labels; real speaker identity / diarization is not silently guessed and remains a planned optional local layer. 會議模式把長對話整理成可審核的逐字稿包。現在可直接吃本機音檔,透過 MLX Whisper adapter 轉成分段逐字稿;也仍支援 JSON/JSONL 分段輸入,方便整合與測試。公開安全預設仍是匿名說話者標籤;真實說話者身分 / diarization 不會偷偷猜,仍是之後的本機可選層。 ```bash python3 daemon/meeting_cli.py export-demo --out /tmp/open-dictate-demo python3 daemon/meeting_cli.py transcribe examples/meeting-segments.example.json --out /tmp/open-dictate-meeting python3 daemon/meeting_cli.py transcribe ~/Desktop/meeting.m4a --out /tmp/open-dictate-meeting-audio --language zh # Mixed English/Chinese recordings can use: --language auto ``` Read: [`docs/MEETING.md`](docs/MEETING.md) ## 自我進化詞庫 / Self-evolving Glossary Open Dictate can scan transcripts for possible mishearings and put candidates into a review queue. Accepted candidates update the user's local glossary; rejected candidates stay out. Open Dictate 可以掃描逐字稿裡的疑似誤聽,把候選放進審核佇列。使用者接受後才寫入本機詞庫;拒絕的候選不會進入確定替換表。 ```bash python3 daemon/qa/mishear_detector.py examples/meeting-segments.example.json --json python3 daemon/glossary/cli.py add "阿布西店" "Obsidian" --reason "near canonical term" python3 daemon/glossary/cli.py candidates ``` Rule: **prefer missed corrections over wrong corrections.** 規則:**寧可漏改,不可錯改。** Read: [`docs/SELF-EVOLVING-GLOSSARY.md`](docs/SELF-EVOLVING-GLOSSARY.md) ## 說話者辨識 / Speaker Identity The public-safe default is anonymous speaker labels: 公開安全預設是匿名說話者標籤: ```text SPEAKER_00 SPEAKER_01 ``` Optional local speaker identity is planned. It must store profiles under local user-owned paths and treat voice embeddings as sensitive biometric data. 可選的本機說話者身分功能在路線圖中。它必須把資料存在使用者本機路徑,並把聲音 embedding 當成敏感生物特徵資料處理。 Read: [`docs/SPEAKER-ID.md`](docs/SPEAKER-ID.md) ## 架構 / Architecture ```text [Hotkey] → OpenDictate.app → /tmp/open-dictate.sock → dictated.py (MLX keep-warm) → deterministic lexicon correction → text insertion → ~/.open-dictate/dictation-log/ [Meeting audio or JSON/JSONL] → meeting_cli.py → local MLX Whisper audio ASR if needed → deterministic correction → anonymous speaker labels → QA flags → Markdown / JSONL / SRT / VTT ``` | Component | Path | |---|---| | Swift shell | `OpenDictate/` | | Dictation daemon | `daemon/dictated.py` | | Meeting CLI | `daemon/meeting_cli.py` | | QA scanner | `daemon/qa/mishear_detector.py` | | Review queue | `daemon/glossary/` | | Speaker layer | `daemon/speaker/` | | Starter glossaries | `vendor/tools/td-subtitle/glossaries/` | | Lexicon engine | `vendor/tools/muse-lexicon/muse_lexicon.py` | | Dictation log | `~/.open-dictate/dictation-log/` | Architecture and protocol: [`IO-CONTRACT.md`](IO-CONTRACT.md) ## 設計原則 / Design Rules 1. Deterministic replacement only: glossary pairs may correct words; the system must not rewrite the sentence. 2. Prefer missed corrections over wrong corrections. 3. Traditional Chinese normalization is allowed; numeric meaning is not changed automatically. 4. Audio stays local. 5. UI must not steal focus from the current text field. 6. Speaker profiles and voice embeddings are local-only sensitive data. 7. Self-evolution is review-first, not silent self-mutation. ## 建置與測試 / Build and Test ```bash ./build.sh python3 -m unittest discover tests python3 scripts/golden-bench.py --skip-daemon python3 scripts/public-safety-scan.py ./scripts/smoke-test.sh ``` The smoke test builds the app, runs deterministic tests, exports a meeting demo, checks helper scripts, and verifies the signed app bundle. ## 路線圖 / Roadmap - [x] Public-safe dictation seed. - [x] Bilingual README and privacy model. - [x] Meeting package MVP for JSON/JSONL segments. - [x] QA flags for possible mishearings and numbers. - [x] Review-first glossary queue CLI. - [x] Anonymous speaker label layer. - [x] Local audio ASR backend for Meeting Mode. - [ ] Optional local speaker profile enrollment. - [ ] Menubar UI for reviewing glossary candidates. - [ ] Import/export for user-owned glossary packages. - [ ] Remove runtime dependency on the source clone, developer Python, Swift, and Xcode Command Line Tools. - [ ] Ship a self-contained beta `.app` with in-app permission onboarding, model download, diagnostics, and uninstall. - [ ] Release a Developer ID signed and Apple-notarized DMG with clean-Mac installation tests. - [ ] Add signed in-app updates, rollback, and Stable/Beta release channels. 完整的 macOS 安裝與發布分階段規劃:[`docs/MACOS-DISTRIBUTION-ROADMAP.md`](docs/MACOS-DISTRIBUTION-ROADMAP.md) Detailed macOS packaging and distribution plan: [`docs/MACOS-DISTRIBUTION-ROADMAP.md`](docs/MACOS-DISTRIBUTION-ROADMAP.md) ## Status Open Dictate is an early public seed extracted from a private daily-use tool, then rebuilt as a public-safe no-history repository. The core dictation path works on Apple Silicon macOS. Meeting, QA, and speaker layers are intentionally conservative MVPs. Open Dictate 目前是早期 public seed,從私人日用工具抽出後重新整理成公開安全、無歷史包袱的 repo。核心語音輸入路徑已可在 Apple Silicon macOS 上使用;會議、QA、說話者層目前是保守 MVP。 Issues and pull requests are welcome. ## License MIT. See [`LICENSE`](LICENSE). --- # Open Dictate Integration Contract v1.0 This file is the source of truth for the Open Dictate app, daemon, glossary engine, and tests. ## System Overview ```text [OpenDictate.app] -- wav path --> [daemon/dictated.py] --> mlx_whisper ^ | deterministic correction via muse_lexicon | v +--------- corrected text ---- log: ~/.open-dictate/dictation-log/ ``` ## Paths | Item | Default Path | Notes | |---|---|---| | Repository | any clone path | `install.sh` resolves paths dynamically | | Swift app | `OpenDictate/` | Swift Package Manager executable | | Daemon | `daemon/dictated.py` | Long-running keep-warm process | | Glossary root | `vendor/` | Override with `OPEN_DICTATE_LEXICON_ROOT` | | Dictation log | `~/.open-dictate/dictation-log/YYYY-MM-DD.jsonl` | Local only; never commit logs | | Socket | `/tmp/open-dictate.sock` | Unix domain socket, newline-delimited JSON | | Python venv | `.venv-dictate/` | Created by standalone install | ## Socket Protocol Requests: ```json {"cmd": "transcribe", "wav": "/tmp/open-dictate-rec-20260705-121501.wav", "punct": "smart_zh"} {"cmd": "ping"} {"cmd": "reload_lexicon"} {"cmd": "add_pair", "wrong": "誤聽", "right": "正確", "source": "dictate-ui"} {"cmd": "stats"} ``` Responses: ```json {"ok": true, "text": "校正後文字", "raw": "whisper 原始輸出", "changes": [["誤聽", "正確"]], "punct": "smart_zh", "asr_ms": 210, "total_ms": 260} {"ok": true, "pong": true, "model": "mlx-community/whisper-large-v3-turbo", "warm": true, "version": "0.6.1", "punct_llm": {"enabled": true, "loading": false, "ready_age_s": 12.3, "last_load_s": 8.9, "last_error": null}} {"ok": false, "error": "no_speech"} ``` Error codes: `no_speech`, `file_not_found`, `asr_failed`, `bad_request`, `unknown_cmd`, `add_pair_failed`. When the repetition-loop guard removes a runaway tail, the local JSONL log entry carries an extra `dehall` field describing what was removed. It never appears on the wire response; the socket contract is unchanged. ⚠️ **Timeout contract**: the shell's `transcribeTimeout()` allows `max(15, audioSeconds/6 + llmHeadroom)` and the daemon's punctuation budget caps at `PUNCT_LLM_TIMEOUT_CAP_S`. These two constants are one contract. Raising the cap without raising the shell headroom produces a false "daemon offline" state while the daemon actually completes the work in the background. Since daemon 0.6.0 the daemon also computes the shell deadline itself (formula above minus a 2 s margin) and never starts or extends an LLM punctuation request past it. `punct` may also be `llm_zh_partial` (daemon ≥0.5.4): long dictation is punctuated in chunks and only some chunks passed; the rest used `smart_zh`. Responses may carry stage timings `pre_asr_ms`, `load_ms`, `lex_ms`, `punct_ms`; `ping` may carry `punct_llm` (warm-keeper status, daemon ≥0.6.0). Readers ignore unknown fields. ## Audio Format - 16 kHz, mono, PCM16 WAV. - Recordings shorter than 0.5 seconds are ignored. - Temporary wav files are written under `/tmp/` and deleted after processing. ## Glossary Schema ```json { "_meta": {"version": "0.1.0", "description": "starter glossary"}, "replacements": {"誤聽": "正確"}, "_review_flagged": {}, "_canonical": ["正確專名"] } ``` Only `replacements` are applied automatically. Ambiguous terms should be flagged for human review instead of being corrected blindly. ## Correction Rules 1. Glossary replacements are deterministic. 2. The system should not rewrite or summarize the sentence. 3. Traditional Chinese normalization and punctuation formatting are allowed. 4. Numbers are not semantically changed automatically. 5. Optional local LLM punctuation must pass the content gate: after punctuation is removed, output content must be reachable from input content using only authorized glossary pairs. ## Quality Gates - Warm transcription of short utterances should target sub-second daemon latency on Apple Silicon. - `smart_zh` must be deterministic and idempotent. - No-rewrite gates must reject inserted, deleted, or changed non-punctuation characters unless they match an authorized pair. - Public fixtures must not contain real user dictation, private names, or private project details. ## Daemon 0.6.0 robustness notes All changes are daemon-internal; the wire protocol stays 1.0. - **LLM warm keeper.** A client timeout while Ollama is loading the punctuation model cancels the load (`context canceled`), so a cold model never becomes warm. A background thread now loads and warms the model with a long timeout and re-checks `/api/ps` every `PUNCT_WARM_INTERVAL_S` (120 s). If the model is not resident, a request falls back to `smart_zh` immediately instead of waiting out a budget. Disable with `_PUNCT_WARM=0`. Cost: the model stays resident. - **Contextual table filtering.** Only contextual pairs whose wrong side occurs in the current text are sent to the LLM and to the reachability gate. Pairs that do not occur cannot be applied, so the gate result is unchanged while the prompt stays small. - **ASR decode guards.** `sample_len = min(224, ceil(seconds × 12) + 32)`; temperature fallback is capped at 0.4 (`(0, 0.4)` below 4 s). High-temperature retries on silence produced loops that took 15–41 s for a one-second accidental recording. - **False-trigger gate.** Recordings under 4 s whose 30 ms-frame dynamic range is below 9 dB and voiced ratio below 0.08 return `no_speech` without running ASR. - **Known subtitle-credit hallucinations** (e.g. `MING PAO CANADA`, `字幕由 Amara.org 社群提供`, `詞曲 李宗盛`) return `no_speech` when they are the whole output; unambiguous credits stuck to the end of real text are trimmed. Output made only of punctuation is `no_speech`. - **Whisper prompt budget.** mlx-whisper keeps only the last 223 prompt tokens. The prompt is now built within 200 tokens with the style sentence at the end; recordings under 4 s use a terms-only prompt. Optional `_PROMPT_CORE_TERMS` (`、`-separated) are always placed first. Small-form punctuation (`﹐﹖﹗`) is normalized. - **Sliver windows.** 30–33 s recordings are split at the quietest point near the middle via `clip_timestamps`, avoiding a sub-second final window that loops. - **Silence trimming** of clear leading/trailing silence (≥0.6 s, only when a quiet floor exists), and a **prompt-swap retry** when a voiced recording ≥4 s yields implausibly little text. - **TAIL_LOOP backtracking fix.** A run of ~40 separators followed by one character made the tail-loop regex backtrack for more than 20 s, blocking the single-threaded daemon. Separator runs are now collapsed before matching (cut positions map back to the original text). - Local JSONL log entries may carry diagnostics: `rms`, `dr`, `voiced`, `asr_temp_max`, `asr_retry`, `raw_first`, `trim_head_s`, `trim_tail_s`; `dehall` may be `speech_gate` or `known_hallucination …`. `no_speech` is a normal outcome, not an error. --- # Open Dictate Setup ## Requirements - Apple Silicon Mac. - macOS 14 or later. - Xcode command line tools. - Python 3.11+ recommended. ## Install ```bash git clone https://github.com/frank890417/open-dictate.git cd open-dictate ./install.sh ``` The installer creates `.venv-dictate`, installs Python dependencies, builds `OpenDictate.app`, installs it to `/Applications/OpenDictate.app`, generates launchd plists with your current clone path, and starts the daemon/app. The first run also downloads and warms the Whisper model. This can take about two minutes depending on the network and disk cache. The installer displays progress for up to 180 seconds and only succeeds after a real daemon ping reports ready; a socket file alone is not considered success. 第一次執行也會下載並載入 Whisper 模型,依網路與磁碟快取狀況可能約需兩分鐘。安裝器最多顯示 180 秒進度,並以實際 daemon ping 為成功條件;只有 socket 檔案不算完成。 If a slow connection needs more time, set `OPEN_DICTATE_WARM_TIMEOUT=300 ./install.sh`. ## Permissions Enable OpenDictate in macOS Privacy & Security: - Microphone - Accessibility - Input Monitoring If you rebuild the app with ad-hoc signing, macOS may require toggling permissions off/on. ### Hotkey conflicts Open Dictate defaults to `fn` push-to-talk. In System Settings → Keyboard: - Set “Press fn/🌐 key to” to “Do Nothing”. - Disable or move Apple Dictation's shortcut if it also uses `fn`. - Quit or reconfigure other dictation apps that capture `fn`. If the hotkey still does not work, re-check Input Monitoring and Accessibility, then quit and reopen OpenDictate. ## Optional External Glossary Root By default Open Dictate uses `vendor/tools/...` starter glossaries. Advanced users may set: ```bash OPEN_DICTATE_LEXICON_ROOT=/path/to/compatible/root ./install.sh --developer ``` The root must contain: ```text tools/muse-lexicon/muse_lexicon.py tools/td-subtitle/glossaries/*.json ``` ## Manual Commands ```bash ./build.sh python3 daemon/dictate_cli.py ping python3 daemon/dictate_cli.py stats python3 scripts/golden-bench.py --skip-daemon ``` ## Diagnose Run the read-only doctor after moving the clone, after a macOS update, or when the menu-bar app cannot reach the daemon: ```bash ./scripts/doctor.sh ``` It checks the installed app, code signature, plist syntax, launchd jobs, clone paths, Unix socket, and a real ping. It only prints reminders for Microphone, Accessibility, and Input Monitoring because it does not inspect or modify the protected macOS TCC database. ## Uninstall ```bash ./uninstall.sh ``` This removes `/Applications/OpenDictate.app` and both LaunchAgents, while preserving `~/.open-dictate` (personal glossaries, transcripts, and logs). ```bash ./uninstall.sh --purge-data ``` For safety, `--purge-data` only prints the local data path and asks you to move it to Trash manually. The script never destroys user data automatically. --- # 隱私模型 / Privacy Model Open Dictate is local-first. Audio, transcripts, logs, review queues, and speaker profiles should stay on the user's Mac unless the user explicitly exports them. Open Dictate 是本地優先工具。音檔、逐字稿、日誌、審核佇列、說話者資料預設都留在使用者自己的 Mac;除非使用者明確匯出,專案不應把它們送出本機。 ## 資料分類 / Data Classes | 類型 | 預設位置 | 是否可進 repo | |---|---|---| | Dictation logs / 語音輸入日誌 | `~/.open-dictate/dictation-log/` | No | | Meeting transcripts / 會議逐字稿 | `~/.open-dictate/meetings/` | No | | Review queue / 誤聽候選審核 | `~/.open-dictate/review-queue/` | No | | Personal glossary / 個人詞庫 | `~/.open-dictate/glossaries/` | No by default | | Speaker profiles / 說話者資料 | `~/.open-dictate/speakers/` | Never | | Public fixtures / 公開測試資料 | `fixtures/` or `examples/` | Yes, fictional or public-domain only | ## Speaker Profiles Are Sensitive Speaker embeddings and voiceprints are biometric data. Treat them like secrets: - Keep them local. - Do not commit them. - Do not paste them into GitHub issues. - Enroll speakers only with consent. - Prefer anonymous `SPEAKER_00`, `SPEAKER_01` labels when sharing transcripts. 聲紋與說話者 embedding 屬於生物特徵資料。請把它們當成敏感資料處理:只放本機、不進 Git、不貼到 issue、取得同意才建檔;要分享逐字稿時,優先使用匿名說話者標籤。 ## Self-evolving Glossary Open Dictate may suggest possible mishearings, but suggestions should enter a review queue first. Accepted pairs can update the user's local glossary; rejected pairs are remembered so the same weak guess is not repeated endlessly. Open Dictate 可以自動提出疑似誤聽,但候選應先進審核佇列。使用者接受後才寫入本機詞庫;拒絕的候選會被記住,避免同一個弱猜測反覆出現。 ## Public Contributions Before opening a pull request, run: ```bash python3 scripts/public-safety-scan.py python3 -m unittest discover tests ``` Do not include real audio, real transcripts, real names, private logs, or speaker profiles in PRs. --- # 會議模式 / Meeting Mode Meeting Mode turns longer recordings or pre-transcribed segments into reviewable transcript packages. 會議模式用來把長音檔或已轉好的分段逐字稿整理成可審核的逐字稿包。 ## Current Public Seed The current public seed accepts two input types: 1. **Local audio files** (`.wav`, `.m4a`, `.mp3`, `.flac`, `.aiff`) through a local MLX Whisper adapter. 2. **Pre-transcribed JSON/JSONL segments** for integrations, fixtures, and testing. 目前公開版支援兩種輸入: 1. **本機音檔**(`.wav`, `.m4a`, `.mp3`, `.flac`, `.aiff`),透過本地 MLX Whisper adapter 轉錄。 2. **已轉好的 JSON/JSONL 分段逐字稿**,給整合、fixture 與測試使用。 ```bash python3 daemon/meeting_cli.py export-demo --out /tmp/open-dictate-demo python3 daemon/meeting_cli.py transcribe examples/meeting-segments.example.json --out /tmp/open-dictate-meeting python3 daemon/meeting_cli.py transcribe ~/Desktop/meeting.m4a --out /tmp/open-dictate-meeting-audio --language zh ``` For mixed English/Chinese meetings, use `--language auto`. To override the ASR model, use `--model mlx-community/whisper-large-v3-turbo` or another compatible MLX Whisper repo. 中英混雜會議可用 `--language auto`。要指定 ASR 模型可加 `--model mlx-community/whisper-large-v3-turbo` 或其他相容的 MLX Whisper repo。 Output: ```text transcript.jsonl transcript.md transcript.srt transcript.vtt meeting-result.json ``` ## Pipeline ```text Audio file → mlx_whisper.transcribe() locally → segment contract → deterministic glossary correction → anonymous speaker labels → QA flags → Markdown / JSONL / SRT / VTT ``` ```text JSON/JSONL segments → segment contract → deterministic glossary correction → anonymous speaker labels → QA flags → Markdown / JSONL / SRT / VTT ``` ## Segment Input Schema ```json { "segments": [ { "start": 0.0, "end": 3.2, "speaker": "alice-local-label", "raw": "今天我們測試 Open Dictate" } ] } ``` Speaker labels are normalized to anonymous labels by default: ```text SPEAKER_00 SPEAKER_01 ``` Audio input currently receives `SPEAKER_00` unless an upstream segment already includes labels. Real speaker identity and cross-meeting voiceprint learning are planned as a separate optional local layer, because voice embeddings are sensitive biometric data. 音檔輸入目前預設標成 `SPEAKER_00`,除非上游分段本身已有 speaker label。真實說話者身分與跨會議聲紋學習會拆成另外一個本機可選層,因為聲紋 embedding 是敏感生物特徵資料。 ## Safety Rule Meeting Mode may flag uncertain spans, but it does not silently rewrite names, numbers, dates, or money-like values. 會議模式可以標記不確定片段,但不會靜默改寫人名、數字、日期或金額類資訊。 --- # 自我進化詞庫 / Self-evolving Glossary Open Dictate's glossary loop is review-first: ```text transcript → QA scanner finds possible mishearings → candidate enters review queue → user accepts / edits / rejects → accepted pair updates local glossary → future transcription improves ``` Open Dictate 的詞庫成長採審核優先:系統可以找候選,但不應把猜測直接寫進確定替換表。 ## Commands ```bash python3 daemon/qa/mishear_detector.py examples/meeting-segments.example.json --json python3 daemon/glossary/cli.py add "阿布西店" "Obsidian" --reason "near canonical term" python3 daemon/glossary/cli.py candidates python3 daemon/glossary/cli.py accept cand_xxxxx python3 daemon/glossary/cli.py reject cand_xxxxx --reason "valid phrase" python3 daemon/glossary/cli.py undo cand_xxxxx ``` ## Buckets - `replacements`: safe deterministic wrong→right pairs. - `_contextual`: risky pairs that need context. - `_review_queue`: pending suggestions. - `_history`: audit trail. ## Rule Prefer missed corrections over wrong corrections. A bad glossary pair can poison every future transcript. 寧可漏改,不可錯改。一個錯誤詞庫 pair 會污染之後每一次轉錄。 --- # 說話者辨識 / Speaker Identity Open Dictate separates two concerns: 1. **Anonymous speaker labels**: `SPEAKER_00`, `SPEAKER_01`. 2. **Optional local speaker identity**: user-owned speaker profiles stored outside the repo. Open Dictate 把兩件事分開: 1. **匿名說話者標籤**:`SPEAKER_00`、`SPEAKER_01`。 2. **可選的本機說話者身分**:由使用者自己建立,存放在 repo 外。 ## Default: Anonymous The public-safe default is anonymous labeling. This is suitable for sharing examples, tests, and public transcripts. 公開安全預設是匿名標籤,適合分享範例、測試與公開逐字稿。 ## Optional: Local Profiles Future local speaker profiles should live under: ```text ~/.open-dictate/speakers/ ``` They must not be committed. They may include biometric embeddings, so treat them as sensitive data. ## Confidence Rule If the system is not confident, it should say `unknown`, not guess a name. 不確定時輸出 unknown,不硬猜人名。 --- # macOS Distribution Roadmap 目標:讓一般使用者從 GitHub Releases 下載 Open Dictate,拖入 Applications、完成系統權限引導後即可使用;不需要 Git、Python、Swift、Xcode Command Line Tools 或保留原始碼 clone。 ## 發布原則 - Open Dictate 公開 repo 是通用產品與發布流程的唯一 SSOT。 - Muse Dictate 只保留私人設定與 adapter,使用同一套公開核心及打包管線。 - 語音、逐字稿、個人詞庫、review queue 與 speaker profile 只存在本機 Application Support,不進 App bundle 或 Git history。 - App 必須能在乾淨的受支援 macOS 帳號安裝、更新、移除與回滾。 ## Phase 0 — Release architecture and contracts - [ ] 將品牌、bundle ID、socket、LaunchAgent label、資料路徑與詞庫 provider 抽成 `ProductConfig`。 - [ ] 將 IO contract 做成 machine-readable schema,並讓公開版與私人 overlay 共用 contract tests。 - [ ] 決定最低 macOS 版本、Apple Silicon 支援矩陣與模型相容政策。 - [ ] 將所有 runtime data 移至 `~/Library/Application Support/OpenDictate/`,logs 移至 `~/Library/Logs/OpenDictate/`。 - [ ] 移除對 repo clone 路徑、開發者 venv 與外部 vendor 目錄的執行期依賴。 **完成條件:** 移動或刪除原始碼 clone 後,已安裝的 App 仍可啟動、聽寫、重啟 daemon 並讀取自己的詞庫。 ## Phase 1 — Self-contained beta app - [ ] 把 Python runtime、daemon、MLX/OpenCC 相依套件、starter glossary 與 helper tools 放入 App bundle 或版本化的 Application Support runtime。 - [ ] 評估 `python-build-standalone` 作為可重定位的 arm64 Python;鎖定 wheels、授權清單與雜湊,不要求使用者安裝 Homebrew 或 Xcode。 - [ ] 由 App 監督 embedded daemon,並以 `SMAppService` 管理登入啟動;使用固定、可升級的 bundle 內路徑,不依賴 shell installer 或外部 LaunchAgent。 - [ ] 首次啟動加入 welcome flow:系統需求檢查、麥克風/輔助使用/輸入監控權限、熱鍵測試、模型準備進度與第一次試聽寫。 - [ ] 將大模型與 App 分離:首次使用時明示大小、儲存位置與下載進度,支援續傳、取消、重試、固定 revision、checksum 驗證和刪除模型。 - [ ] 將既有 `~/.open-dictate` 資料做成一次性、可回復的 migration,不遺失詞庫或設定。 - [ ] 加入 App 內 doctor、重啟服務、重設權限說明、匯出診斷資料與完整解除安裝。 - [ ] 產出可供測試的 `.dmg`;beta 階段仍可手動發布,但不得要求使用者執行 Terminal 指令。 **完成條件:** 在乾淨的 Apple Silicon Mac 上,只靠 Finder 與 App 內引導,在 10 分鐘內完成安裝及第一次聽寫;移除下載來源資料後仍可離線使用。 ## Phase 2 — Signed and notarized public release - [ ] 申請並設定 Developer ID Application 憑證與 hardened runtime。 - [ ] 補齊 entitlements;由內而外簽署 Python/MLX nested executables、dylibs、helper 與 App,不能以 `codesign --deep` 取代正確的簽署順序。 - [ ] 完成 Apple notarization、stapling,以及 `codesign --strict`、`spctl`、`stapler validate` 驗證。 - [ ] 發布簽章、公證過的 DMG;附 SHA-256 checksum、版本資訊、隱私說明與支援矩陣。 - [ ] 在乾淨 macOS VM/實機執行 Gatekeeper、安裝、權限、重開機、聽寫、doctor、uninstall 與 rollback 測試。 - [ ] CI release workflow 只能從受保護 tag 產生 artifact;公開安全掃描與 secret/history/blob 掃描必須先通過。 **完成條件:** 使用者雙擊 DMG、拖入 Applications 後可正常開啟,Gatekeeper 不顯示未識別開發者警告,重開機後服務能恢復。 ## Phase 3 — Updates, rollback, and release operations - [ ] 導入 Sparkle 2 或等價的簽章更新框架,使用 EdDSA appcast;更新 metadata 與 binary 分離簽章。 - [ ] runtime、模型與使用者資料分開版本化;App 更新不得覆蓋個人詞庫、紀錄或 speaker profile。 - [ ] 支援更新前健康檢查、失敗自動退回前一版、保留一個已知可用 runtime。 - [ ] 建立 Stable/Beta 更新頻道與版本支援政策;security/protocol 修正可要求立即更新。 - [ ] 每個 Open Dictate release 自動觸發 Muse Dictate overlay 相容性測試與版本升級提醒。 **完成條件:** 能從前一個正式版本在 App 內升級,保留所有使用者資料;刻意注入失敗更新時可自動退回且仍能聽寫。 ## Phase 4 — Optional distribution channels - [ ] 評估 Homebrew Cask,作為開發者與進階使用者的第二安裝管道。 - [ ] 等 sandbox、background service、模型下載與 Accessibility/Input Monitoring 權限限制確認可接受後,再評估 Mac App Store。 - [ ] 若增加 Intel 支援,必須獨立驗證 ASR backend、效能與 artifact;不能只做 universal binary 就宣稱支援。 ## 建議的 artifact 邊界 ```text OpenDictate.app ├── Swift UI and onboarding ├── signed background service ├── private Python runtime and native libraries ├── daemon and deterministic lexicon engine └── public starter glossary ~/Library/Application Support/OpenDictate/ ├── models/ ├── glossary/ ├── review-queue/ ├── speaker-profiles/ └── runtime-state/ ``` 模型預設不塞進 DMG:這能讓 App 下載較小、模型獨立升級,也讓使用者在首次啟動時清楚同意磁碟用量。若未來需要完全離線安裝,可另外提供含模型的大型 offline artifact,不與一般版本混在一起。 第一個正式安裝格式採 DMG,不採 PKG。現階段沒有需要 root 權限、system extension 或系統層安裝位置的元件;DMG 的安裝與移除較透明。若未來真的增加系統級元件,再重新評估 PKG。 ## 暫不採用 - 不把現有 `install.sh` 直接包成 `.pkg`:它仍依賴開發工具與 clone 路徑,沒有解決可攜性。 - 不用 ad-hoc 簽章作公開發布:重新建置會破壞 TCC 權限身分,也無法提供正常 Gatekeeper 體驗。 - 不把模型、個人詞庫或 runtime data 永久寫進可覆蓋的 App bundle。 - 不先追求 Mac App Store;目前的全域熱鍵、Accessibility、Input Monitoring 與 background service 需要先做 sandbox 可行性驗證。 ## 技術參考 - [Apple:發佈前公證 macOS 軟體](https://developer.apple.com/documentation/security/notarizing-macos-software-before-distribution) - [Apple:建立 macOS 發布簽章](https://developer.apple.com/documentation/xcode/creating-distribution-signed-code-for-the-mac/) - [Apple:SMAppService](https://developer.apple.com/documentation/servicemanagement/smappservice) - [Apple:更新 App package installer 至新的 Service Management API](https://developer.apple.com/documentation/servicemanagement/updating-your-app-package-installer-to-use-the-new-service-management-api) - [Sparkle:發布與更新簽章](https://sparkle-project.org/documentation/publishing/) --- # MCP roadmap The MCP surface is a local automation boundary for Open Dictate. It is not a remote transcription service and does not grant an agent ambient access to the microphone, recordings, transcripts, personal glossary, or speaker profiles. ## Product principles - Start with a local `stdio` server launched by an MCP host. Do not listen on a network port in the first release. - Default to read-only discovery. Every mutation or recording action requires explicit, narrowly scoped user consent at invocation time. - Reuse the versioned daemon protocol and `ProductConfig`; MCP must not become a second source of truth for transcription or storage paths. - Return the minimum necessary data. Prefer opaque IDs and summaries over raw audio paths, full transcripts, or personal dictionary contents. - Keep private memory and identity implementations behind adapters. The public server defines capabilities and consent boundaries only. ## Phase M0 — Contract and threat model - Publish a capability manifest and JSON Schemas for MCP inputs and outputs. - Define stable identifiers for recordings, transcript jobs, glossary proposals, and diagnostics without exposing absolute filesystem paths. - Document trust boundaries for the MCP host, local daemon, external lexicon adapter, filesystem, and optional model downloader. - Add adversarial tests for path traversal, symlinks, oversized payloads, newline injection, prompt injection in transcripts, and concurrent mutations. **Exit gate:** every proposed capability has an owner, data classification, consent rule, audit event, size limit, timeout, and redaction behavior. ## Phase M1 — Local read-only server Implement a dependency-pinned local `stdio` server after choosing an SDK with a maintained security policy. Initial resources: - `opendictate://status`: app, daemon, model, protocol, and health summary. - `opendictate://capabilities`: supported commands and optional adapters. - `opendictate://glossary/summary`: counts and schema version, not term contents. - `opendictate://diagnostics/recent`: redacted local diagnostic summaries. Initial tools are non-recording and non-mutating: - `dictate_health_check` - `dictate_list_capabilities` - `dictate_validate_audio` for an explicitly supplied file handle or approved ID Prompts may offer meeting-preparation and glossary-review templates, but prompt arguments are untrusted content and never authorize a tool call. **Exit gate:** server works with the App stopped where appropriate, has no network listener, reveals no private paths, and passes host interoperability and malformed-message tests. ## Phase M2 — Consent-gated actions Add narrowly scoped tools only after the host can display informed consent: - `dictate_transcribe_audio`: process one user-selected audio object. - `dictate_propose_glossary_pair`: create a review item; never auto-accept it. - `dictate_export_transcript`: export one named job to an approved destination. - `dictate_reload_lexicon`: reload after an external, user-approved change. Microphone capture remains an App interaction. If a future MCP tool can start recording, each start requires foreground confirmation, shows a persistent recording indicator, enforces a time limit, and provides an immediate stop tool. There is no background or blanket recording consent. **Exit gate:** cancellations stop work, mutation tools are idempotent, every change has a local audit event, and denial leaves no partial output. ## Phase M3 — Private memory and connection adapters Define public interfaces for optional memory lookup, glossary suggestion, meeting context, and speaker-label providers. Private products may implement them out of tree. The public MCP server receives only the minimal result needed for the active invocation and must operate correctly when no adapter is present. Adapters declare capability and schema versions. They cannot add undeclared MCP tools dynamically, bypass consent, weaken redaction, or write to public core storage. Compatibility is checked using the same overlay lock and release gates. ## Tool annotations and safety Each tool declares `readOnlyHint`, `destructiveHint`, `idempotentHint`, and `openWorldHint` where the selected MCP SDK supports them. Treat annotations as UI and planning hints, not access control. The server independently enforces: - allowlisted operations and bounded input sizes; - canonical paths under approved roots with symlink checks; - socket ownership and `0600` permissions; - explicit destination approval for exports; - redaction of paths, transcript text, glossary terms, model prompts, and tokens; - local-only audit records with retention controls; - no shell evaluation and no command strings supplied by clients. Resources containing transcript text or glossary entries are opt-in and scoped to an explicit object ID. They are never exposed as broad enumerable resources. ## Evaluation matrix MCP releases require protocol conformance tests, two supported host smoke tests, read-only snapshot tests, consent denial and cancellation tests, mutation idempotency tests, daemon unavailable/restart tests, private-adapter absent and incompatible tests, data-leak canaries, and latency/timeout budgets. A red-team fixture places instructions inside a transcript and verifies that they remain data rather than becoming tool authorization. No MCP package is shipped until its dependencies are locked, licenses audited, and the stdio process can be disabled or removed without affecting ordinary Open Dictate use. --- # Contributing to Open Dictate Thanks for helping improve Open Dictate. 謝謝你一起改善 Open Dictate。 ## Ground Rules - Keep examples fictional or public-domain. - Do not commit real audio, real transcripts, dictation logs, private paths, personal glossaries, or speaker profiles. - Do not paste private transcripts or voiceprints into GitHub issues. - Prefer deterministic correction and review queues over silent rewriting. - Run the safety and test commands before opening a pull request. ## Before Pull Request ```bash python3 scripts/public-safety-scan.py python3 -m unittest discover tests python3 scripts/golden-bench.py --skip-daemon ``` If you change Swift code, also run: ```bash ./build.sh ``` ## Privacy Read [`docs/PRIVACY.md`](docs/PRIVACY.md). Speaker profiles and voice embeddings are sensitive biometric data. Public fixtures must use fake examples only. --- # Security Policy Open Dictate processes microphone audio locally. Please do not attach private recordings, transcripts, dictation logs, or screenshots containing sensitive text to public issues. If you find a vulnerability, open a GitHub security advisory or contact the maintainer privately through GitHub. Include reproduction steps using fictional test data whenever possible.