erm: Bring Voice, Media and Operational Knowledge to Your Desktop
Keep a small command from becoming another context switch
Pause the music. Find a track. Ask a question about an indexed document. These are modest interactions, but they repeatedly pull attention away from the task on screen. erm brings them into a supervised Erlang desktop workspace with voice input, media control, optional local model assistance and speech output.
The clearest first audience is a Linux developer or operator willing to run the local audio and GTK dependencies. Start with a few observable commands on your own workstation. That is enough to evaluate whether the interaction helps without treating the whole desktop as an autonomous agent.
Common commands have a direct path
The voice-intent module recognises play, pause, stop, next, previous, volume, track selection and player visibility before asking a model to classify an unknown phrase. Volume input is bounded. Known negated commands are refused. The media adapter sends typed actions to the existing MPV process owner and uses the playlist state for track selection and progression.
The voice coordinator handles transcript boundaries, command windows, deduplication and asynchronous work. Model planning and media actions have timeouts; they do not run as blocking work inside the coordinator's message handler. The operating benefit is that command state can be inspected and failures can be reported through the same service interface.
Try a small command set such as βBob, pause musicβ, βBob, next songβ and βBob, volume 30β after configuring the trigger. These are pilot utterances, not measured recognition results. Check actual playback and volume after each command, including repetitions and silence around the wake phrase.
Model assistance has a defined action vocabulary
For phrases outside the direct parser, the code requests a bounded JSON intent and validates it against built-in or operator-configured actions. Model output does not choose arbitrary Erlang module/function pairs. Custom callbacks come from trusted configuration.
That is a useful extension boundary for a workstation tool. A maintainer can add a named action with an explicit implementation and test its effect. The callback still carries whatever authority its implementation has; a voice phrase or speaker match should not be treated as approval for payments, secret access or other sensitive operations.
The configured Ollama path can answer a general question or assist with an intent. Local processing depends on the actual configuration and installed models. The code also supports pulling a missing configured model, so setup and downloads are separate from an offline-use claim.
Ask indexed knowledge, with the retrieval boundary visible
The βask ECAIβ route retrieves up to three sources from the configured disk corpus and sends bounded excerpts to the answer model. Missing corpus configuration and empty retrieval have explicit error paths.
This is a concrete connection between erm and ECAI, but it uses
ecai_ollama_rag:retrieve_sources/3. It does not call the corpus/principal
private bridge. Use an appropriate public or non-sensitive corpus for this
pilot. Private voice access needs an explicit authenticated integration with
the private retrieval boundary.
Speech and service health are part of the interface
The TTS service exposes native Piper speech, voice selection, volume controls, repeat, cancellation and diagnostics. Model installation is handled by an asynchronous service with manifest-driven file checks and immutable cache directories. Listening suppression is used around speech output to reduce feedback into the recogniser.
The optional native voice coordinator also contains enrolment and speaker-score handling. The source archive does not include its native capture/Whisper binary implementation, and no audio was tested here. Speaker matching should be evaluated with recordings, other speakers and replay attempts before it is given any security role; the reviewed code does not establish liveness.
In a running configured release, these inspection calls help separate a command problem from a missing model or disconnected speech backend:
erm_voice:status().
erm_voice:healthcheck().
erm_tts:diagnostics().
Readiness is not an end-to-end microphone test. Verify capture, recognition, dispatch, playback and the spoken response as one real interaction.
A desktop that can grow around the operator
Lens provides an optional GTK interface with Nostr feed, media and configured wallet integration components. Its supervision and show operation return useful startup errors, allowing it to be evaluated separately from voice. Wallet adapters, mainnet permission and confirmation behaviour need their own setup and tests. No payment or wallet action is part of the suggested voice pilot.
The first useful result is a workstation that performs a small set of commands reliably and makes its failures visible. Measure accidental triggers, missed commands and time to recover a disconnected backend on your own hardware.
Source basis and next steps
Reviewed modules under apps/erm/src/: erm_voice, erm_voice_boundary,
erm_voice_intent, erm_voice_media, erm_native_voice, erm_voice_health,
erm_tts, erm_model_pull, erm_lens, erm_lens_wallet and erm_sup.
Isolated voice/TTS fixtures are present under apps/erm/test/; fake services
and ports in those fixtures are not evidence of live hardware operation.
