Sonora

Sonora is a lightweight, non-autoregressive text-to-speech (TTS) model line built for real-time, on-device, and resource-constrained environments — the voice engine of Project Prosodia. Architecture: Matcha-TTS-family Optimal-Transport Conditional Flow Matching (OT-CFM) acoustic model + HiFi-GAN vocoder, with a growing set of expressive conditioning channels (valence / arousal-energy / tension via FiLM) aimed at directable, story-driven narration.

This repository is the model registry: each top-level directory is one promoted, audited milestone artifact with its own README, eval report, and provenance.

Registry index

Current trunk: derisk-energy-24k (acoustic) + vocoder-24k-hifigan (vocoder) — the 24 kHz multi-speaker line. Next up: the v1.1 VAT continuation (story-driven valence corpus), warm-starting from vat3-24k. Machine-readable lineage: registry.json.

Directory Training pass What it is Status
baseline-ljspeech-22k/ Phase 0 Single-speaker LJSpeech fine-tune, 22.05 kHz, end-to-end graphs (ONNX + TFLite) and the LiteRT split-graph mobile lane. (Renamed from v1-ljspeech on 2026-07-22 — update old deep links.) Published 2026-07-12
vocoder-24k-hifigan/ 24 kHz trunk — vocoder HiFi-GAN vocoder fine-tune, 24 kHz / 80-band — the vocoder pairing for the multi-speaker line. Converged + human-audited (copy-synthesis A/B indistinguishable). Published 2026-07-16
derisk-energy-24k/ Expressive de-risk (north-star §7) First trained expressive channel: energy/arousal FiLM conditioning, multi-speaker (247 spk, LibriTTS-R), 24 kHz. Controllability ρ≈1.0, identity-preserving, WER-safe; includes the VAT-ready LiteRT split-graph export. Published 2026-07-16
vat3-24k/ Full 3-channel VAT (milestone 3) First all-three-channels training (valence/energy/tension FiLM), warm-started from derisk-energy-24k. Mixed verdict: energy PASS, tension near-pass, valence FAIL — a corpus-label limitation, not a training one (see its README). Warm-start seed for the v1.1 continuation. Published 2026-07-22

Directory names are deliberately descriptive, not phase-numbered — artifacts keep their identity even as the roadmap's phase labels evolve; the "Training pass" column carries the roadmap mapping.

Developed by: Artificial Humanity · License: Apache-2.0 · Languages: English

Model line at a glance

  • Phase 0 (baseline-ljspeech-22k) proved the on-device lane: ~18.2M-parameter acoustic model, single forward pass text→PCM, running on phones via TFLite/LiteRT.
  • The 24 kHz multi-speaker line (vocoder-24k-hifigan + derisk-energy-24k) is the current trunk: native 24 kHz, 247 speakers, and validated FiLM conditioning — the de-risk experiment demonstrated a continuous, monotonic, speaker-preserving loudness/ energy control learned from weak labels in ~35 GPU-hours on consumer hardware.
  • Full 3-channel VAT conditioning (valence/arousal/tension) and story-driven expressive training data (directed teacher synthesis + aligned public-domain audiobooks) are in active development; artifacts land here as they pass their gates.

Provenance & audit culture

Every artifact directory documents: training corpus + license wall, warm-start lineage, pre-registered convergence criteria, automated eval-harness results (controllability / identity leakage / intelligibility), and human audition verdicts. Checkpoints are distributed as safetensors (weights-only); full resumable training checkpoints are retained privately.

Companion dataset

The sonora-expressive-registers dataset (CC-BY-4.0) — certified keeps from our directed teacher-synthesis pipeline, with exact intended-VAT labels and full render provenance — is published separately under this organization.

Downloads last month
82
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using artificial-humanity/Sonora 1