Habibi-TTS
Arabic

Licence of the in-house ALG data, and commercial licensing of the Unified model

#1
by lamine-sanfourdz - opened

Hi, and thank you for releasing Habibi and the benchmark — the first open
unified-dialectal Arabic TTS is a significant contribution.

I'm evaluating Habibi-derived models for a commercial product in Algeria,
and I have two questions the paper left open:

(1) Table 1 lists the ALG training data as 64.4h "in-house" (70,970 utterances)
plus 8.6h from the Omnilingual ASR Corpus. §2.2 also mentions "manually
transcribed public speech recordings". What is the source and licence of the
in-house ALG portion? The ALG checkpoint is released under Apache-2.0, but
the paper describes the work as "purely a research project", so I want to be
sure the data permits commercial use of models derived from that checkpoint.

(2) The Unified model (CC-BY-NC-SA-4.0) outperforms the ALG-specialized model
on Algerian dialect accuracy (DMOS 3.90 vs 3.68) and on speaker similarity
(SMOS 4.10 vs 3.78, SIM 0.731 vs 0.306 for ElevenLabs). Is there any path to
a commercial licence for the Unified checkpoint, or a dual-licensing option?

Thanks for the work — the Algerian results in particular are the best I've
seen from any system, open or commercial.

Hi, yes, the ALG-specialized model permits commercial use.
The unified model is of NC-SA because SADA dataset's restriction (SADA is NC-SA).

Sign up or log in to comment