Plume uses Windows UI Automation (UIAutomation TextPattern/ValuePattern) to detect what you are typing. When you stop typing, it reads the last word from the focused text field, runs it through Hunspell for spelling corrections, sends the surrounding context to your chosen LLM for next-word suggestions, and shows the results in a floating transparent overlay. Click a chip to copy or insert.
Floating typing assistant for Windows with voice-to-text dictation. Plume sits as a transparent overlay above your active window, listens via UI Automation (no keyloggers), serves up smart suggestions — Hunspell-powered spelling corrections, AI next-word predictions, translation, text actions — and transcribes your speech via whisper.cpp.
Capture
UI Automation — reads committed text from any text field. No WH_KEYBOARD_LL, no GetAsyncKeyState.
Voice Dictation
Speech-to-text via whisper.cpp (CPU/CUDA/Vulkan). Cuts at natural speech pauses.
Corrections
Hunspell dictionaries with prefix matching — works fully offline.
AI Suggestions
Next-word prediction using a local LLM (llama.cpp) or cloud APIs (OpenAI, Ollama, custom).
Translation
Translate captured text into any language. Rephrase, simplify, change tone.
Privacy
Never records keystrokes. Audio processed locally. Network egress only when you opt into a remote API.
How It Works
Plume uses Windows UI Automation (UIAutomationTextPattern / ValuePattern) to detect what you're typing. When you stop typing, it:
Reads the last word from the focused text field
Runs it through Hunspell for spelling corrections
Sends the surrounding context to your chosen LLM for next-word suggestions
Shows the results in a floating transparent overlay — click a chip to copy or insert
The overlay stays on top of all windows, auto-hides after a configurable idle timeout, and repositions with your active window.
Start typing in any app — suggestions appear automatically
AI Providers
Plume supports four backends. Configure yours in Settings (gear icon on the overlay).
Local (llama.cpp)
Uses a bundled llama-server.exe to run GGUF models locally. The model file is auto-downloaded from HuggingFace on first use.
Preset
Size
Notes
Qwen3-0.6B (Q8_0)
639 MB
Recommended — near-lossless, great for text continuation
Qwen2.5-0.5B (Q8_0)
507 MB
Smallest, lowest RAM
Llama 3.2-1B (IQ3_M)
627 MB
Multilingual
Gemma 3-1B (Q4_K_M)
769 MB
Efficient on CPU
Qwen3.5-2B (Q4_K_M)
1.28 GB
Higher quality
Qwen3.5-4B (Q4_K_M)
2.74 GB
Higher capability
Qwen3.5-9B (Q4_K_M)
5.68 GB
Most capable
Gemma 2-2B (IQ3_M)
1.33 GB
Most capable
Select Custom... to type any GGUF filename. Built-in presets download from verified HuggingFace URLs; custom filenames auto-resolve from unsloth.
Ollama
Connect to a local or remote Ollama instance. Supports any model you've pulled.
Preset
Notes
qwen3:0.6b
Lightweight, fast
qwen3:1.7b
More capable
llama3.2:1b
Multilingual
tinyllama:1.1b
Compact
smollm2:1.7b
Code-aware
phi3:mini
Microsoft's small model
OpenAI-compatible
Connect to OpenAI or any OpenAI-compatible API (together.ai, Groq, etc.). Set your API key and endpoint.
Preset
Input/Output (per 1M tokens)
GPT-5.4 Nano
$0.20 / $1.25
GPT-5.4 Mini
$0.75 / $4.50
GPT-4.1 Mini
$0.40 / $1.60
GPT-4o Mini
$0.15 / $0.60
GPT-5.5
$5.00 / $30.00
GPT-5.5 Pro
$30.00 / $180.00
Custom API
Bring your own endpoint with custom headers (JSON). Useful for self-hosted or proxy setups.
Features
Spelling Corrections
Hunspell parses .dic/.aff files directly — no external process. Corrections use prefix matching (binary search) for fast lookup. Click a correction to insert it. Click the copy icon to copy to clipboard.
AI Next-Word Suggestions
After you stop typing, Plume sends the last ~10 words to the LLM and shows predicted next words as special chips with an "AI" indicator. The suggestion delay is configurable in Settings (default 800ms).
Translation
Toggle translation in Settings. Select a target language, type or capture text, and click the translate button. Results can be copied or inserted directly.
Text Actions
Select an action from the dropdown and apply it to the captured text:
Click the mic button on the overlay toolbar to start real-time voice-to-text. Plume captures your microphone audio, resamples to 16 kHz, and transcribes via whisper.cpp.
Silence-based chunking — audio is accumulated continuously. When a natural speech pause is detected (~350ms of silence), the chunk is transcribed immediately. If you speak without pausing, it forces a cut at the configured max chunk length (default 2s).
Configurable noise threshold — adjust the silence detection level in Settings to ignore background noise.
GPU acceleration — supports CPU, CUDA (NVIDIA), and Vulkan (any GPU with vulkan-1.dll). Falls back to CPU automatically if the GPU backend fails.
Level indicator — three green volume bars on the mic icon show real-time mic level.
Whisper model download — tiny (75 MB, multilingual) by default; choose larger models (base, small) in Settings.
Transcribed text appends directly into the currently focused text field, preserving what you typed manually.
The capture loop is automatically suspended while dictating to avoid clobbering the transcript.
System audio capture — select an output device to capture audio already playing through your speakers (loopback). Both microphone and system audio are captured simultaneously for independent transcription. Enable in Settings under the Output section.
Whisper Model
Size
Language
tiny
75 MB
Multilingual (recommended)
tiny.en
75 MB
English only (fastest)
base
142 MB
Multilingual
small
466 MB
More accurate, slower
Window & Overlay
Transparent, always-on-top, frameless window with rounded corners
Position saved on move/resize, restored on restart
System tray with left-click toggle and right-click menu (Show/Hide, Settings, Quit)
Configuration
Plume stores settings in %APPDATA%\plume\config.json. You can edit this file directly — all fields use #[serde(default)], so missing keys never break deserialization.
Key settings
Key
Default
Description
provider
"local"
local, ollama, openai, or custom
model
"Qwen3-0.6B-Q8_0.gguf"
Model name / GGUF filename
endpoint
"http://127.0.0.1:8080"
API endpoint
port
8080
Local llama-server port
suggestion_count
6
Number of spelling corrections
ai_suggestion_count
3
Number of AI next-word suggestions
ai_suggestion_delay
800
Delay before AI suggestions fire (ms)
idle_timeout
6
Seconds before overlay auto-hides
dictionary.language
"en_US"
Hunspell dictionary code
max_tokens
8192
Max output tokens for LLM requests (summarize, translate, etc.)
audio.model
"ggml-tiny.bin"
Whisper model filename
audio.chunk_secs
2
Max chunk length (seconds) before forced cut
audio.silence_threshold
0.01
RMS silence threshold (0–0.1); lower = more sensitive
audio.compute_backend
"cpu"
Whisper backend: cpu, cuda, or vulkan
audio.language
"auto"
Whisper language (auto for auto-detect, or a language code)
audio.input_device
""
Microphone device ID (empty = system default)
audio.output_device
""
Output device for loopback capture (empty = system default)
audio.enable_output_capture
false
Enable system audio (loopback) capture from output devices
Keyboard Shortcuts
Shortcut
Action
Click chip
Insert suggestion into active field
Copy icon
Copy suggestion to clipboard
Mic button
Toggle voice-to-text dictation
Gear icon
Open Settings
Tray left-click
Toggle overlay visibility
Tray right-click
Menu: Settings, Quit
Building from Source
npm install
npm run tauri dev # Development mode
npm run tauri build # Production build