Sonora
Sonora is a lightweight, non-autoregressive text-to-speech (TTS) model line built for real-time, on-device, and resource-constrained environments — the voice engine of Project Prosodia. Architecture: Matcha-TTS-family Optimal-Transport Conditional Flow Matching (OT-CFM) acoustic model + HiFi-GAN vocoder, with a growing set of expressive conditioning channels (valence / arousal-energy / tension via FiLM) aimed at directable, story-driven narration.
This repository is the model registry: each top-level directory is one promoted, audited milestone artifact with its own README, eval report, and provenance.
Registry index
Current trunk:
derisk-energy-24k(acoustic) +vocoder-24k-hifigan(vocoder) — the 24 kHz multi-speaker line. Next up: the v1.1 VAT continuation (story-driven valence corpus), warm-starting fromvat3-24k. Machine-readable lineage:registry.json.
| Directory | Training pass | What it is | Status |
|---|---|---|---|
baseline-ljspeech-22k/ |
Phase 0 | Single-speaker LJSpeech fine-tune, 22.05 kHz, end-to-end graphs (ONNX + TFLite) and the LiteRT split-graph mobile lane. (Renamed from v1-ljspeech on 2026-07-22 — update old deep links.) |
Published 2026-07-12 |
vocoder-24k-hifigan/ |
24 kHz trunk — vocoder | HiFi-GAN vocoder fine-tune, 24 kHz / 80-band — the vocoder pairing for the multi-speaker line. Converged + human-audited (copy-synthesis A/B indistinguishable). | Published 2026-07-16 |
derisk-energy-24k/ |
Expressive de-risk (north-star §7) | First trained expressive channel: energy/arousal FiLM conditioning, multi-speaker (247 spk, LibriTTS-R), 24 kHz. Controllability ρ≈1.0, identity-preserving, WER-safe; includes the VAT-ready LiteRT split-graph export. | Published 2026-07-16 |
vat3-24k/ |
Full 3-channel VAT (milestone 3) | First all-three-channels training (valence/energy/tension FiLM), warm-started from derisk-energy-24k. Mixed verdict: energy PASS, tension near-pass, valence FAIL — a corpus-label limitation, not a training one (see its README). Warm-start seed for the v1.1 continuation. |
Published 2026-07-22 |
Directory names are deliberately descriptive, not phase-numbered — artifacts keep their identity even as the roadmap's phase labels evolve; the "Training pass" column carries the roadmap mapping.
Developed by: Artificial Humanity · License: Apache-2.0 · Languages: English
Model line at a glance
- Phase 0 (
baseline-ljspeech-22k) proved the on-device lane: ~18.2M-parameter acoustic model, single forward pass text→PCM, running on phones via TFLite/LiteRT. - The 24 kHz multi-speaker line (
vocoder-24k-hifigan+derisk-energy-24k) is the current trunk: native 24 kHz, 247 speakers, and validated FiLM conditioning — the de-risk experiment demonstrated a continuous, monotonic, speaker-preserving loudness/ energy control learned from weak labels in ~35 GPU-hours on consumer hardware. - Full 3-channel VAT conditioning (valence/arousal/tension) and story-driven expressive training data (directed teacher synthesis + aligned public-domain audiobooks) are in active development; artifacts land here as they pass their gates.
Provenance & audit culture
Every artifact directory documents: training corpus + license wall, warm-start lineage, pre-registered convergence criteria, automated eval-harness results (controllability / identity leakage / intelligibility), and human audition verdicts. Checkpoints are distributed as safetensors (weights-only); full resumable training checkpoints are retained privately.
Companion dataset
The sonora-expressive-registers dataset (CC-BY-4.0) — certified keeps from our
directed teacher-synthesis pipeline, with exact intended-VAT labels and full render
provenance — is published separately under this organization.
- Downloads last month
- 82