A transducer rather than an attention decoder, which is a genuinely different way to turn acoustics into tokens. Its declared language set is the widest European one we run, and it is the only engine here that lists Ukrainian alongside Russian.
- Selector
MODEL_TYPE=parakeet
- Architecture
- 24-layer FastConformer encoder, d_model 1024, 8 heads, depthwise-striding subsampling by 8, 128 mel features. LSTM prediction network with 640 hidden units, joint network, and greedy TDT decoding over durations 0 to 4. SentencePiece detokenisation.
- Parameters
- 627,090,606 exact, read from the checkpoint
- Languages
- 25 European languages en, de, es, fr, it, pt, nl, pl, ro, hu, cs, sk, bg, hr, sl, uk, ru, sv, da, fi, nb, el, ca, eu, gl
- Segments
- No segments or timestamps
- Upstream
nvidia/parakeet-tdt-0.6b-v3
- Rate
- ~$0.75 per audio hour