We are FOSS voice, AI, and data specialists. As creators of the HiveMind stack and core contributors to OpenVoiceOS, we build privacy-first, GDPR-compliant voice technology that keeps your data on your own hardware — and we are equally at home turning hard-to-reach public data into clean APIs and curated datasets.
How We Work
The standard engagement is a monthly retainer for ongoing access and support, with discrete deliverables — datasets, trained models, custom plugins, integrations — scoped and billed separately as they are defined. This keeps the relationship predictable on both sides: you have direct access to the team, and each concrete output has its own scope and price.
We also take on standalone project work where the scope is well-defined from the start.
We Voice-Enable Anything
We create voice interfaces for anything. Whatever you want to talk to — a device, an app, a service, or a whole product line — we build the voice layer for it, reusing the expertise and battle-tested stack we’ve built around the Open Voice Operating System and HiveMind. Custom wake words, Text-to-Speech (TTS) and Speech-to-Text (ASR) wiring, intent handling, and bespoke components — assembled from a proven open-source foundation instead of from scratch, and running on your own hardware.
Custom Offline TTS Voice Training
Want a unique voice for your project? We train custom, high-quality, fully-offline Text-to-Speech voices. Bring a dataset for a specific language, or clone a reference voice — the result is a fast, private TTS model that runs locally on your device and slots straight into OpenVoiceOS or Home Assistant. You keep full control of your data.
Curious what our voices sound like? Try the Voices Demo — our Miro & Dii voices synthesize live in your browser, in 20+ languages.
Fine-Tuned Speech Recognition for Your Language & Domain
We fine-tune state-of-the-art speech-to-text models — Zipformer, Parakeet, Whisper, and others — for your specific language, accent, and domain vocabulary, so recognition holds up on the words and conditions that actually matter to you. When the training data doesn’t exist yet, we find it — and synthesize it where appropriate — using our dataset-construction toolchain. This work is backed by our partnership with Alpha Cephei, bringing decades of automatic speech recognition expertise to your project.
Minority & Lusophone Language Speech Tech
We build speech technology for languages that mainstream tools ignore — including Portuguese variants and other Lusophone and minority languages. From phonemization to ASR and TTS, we help under-served language communities get modern, privacy-respecting voice support.
Data Extraction, Clean APIs & Datasets
A great deal of valuable data lives on the public web but is effectively unreachable — locked in unstructured pages, behind anti-bot defenses, exposed only through undocumented endpoints, or in formats no search engine indexes. We get it out and make it usable:
- Resilient scrapers & parsers for sources that resist automation, engineered to keep working as pages and defenses change.
- API reverse-engineering — we map undocumented or private endpoints and re-expose them as clean, typed, documented client libraries.
- One clean API over messy sources — a single consistent interface instead of bespoke scraping for every consumer.
- Dataset construction — we curate the extracted data into structured, versioned, ML-ready datasets with clear provenance, and we can publish open datasets where it makes sense.
It is the same toolchain behind our own client libraries and the datasets we release — and it feeds everything from media-metadata enrichment to training corpora for speech and language models.
Voice-Enabled, Accessible Websites
The same voice stack we ship on devices also runs in the browser. With phoonnx.js, our TTS voices synthesize client-side — no cloud API, no per-request cost, nothing leaving the visitor’s machine. The Voices Demo on this very site is the proof: a static page that speaks.
We build websites and web applications around that capability — sites that can read themselves aloud, voice-first interfaces, and accessible-by-design pages for users who browse by ear. Performance, accessibility, and privacy are the defaults, not extras. We also develop and maintain the official OpenVoiceOS website.