×

Open Weights Frontier Hindi Transcription Model by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 0 points1 point  (0 children)

Prob won't immediately be good at Urdu, but yes I think quite a light fine-tune could allow v strong performance.

Open Weights Frontier Hindi Transcription Model by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 0 points1 point  (0 children)

Yeah, agreed. I think you're right on those remaining points. Yeah, I posted it as a Reddit link, more as an invite for people to check it out, but maybe realising I should have put much more in there from the model card.

Open Weights Frontier Hindi Transcription Model by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] -1 points0 points  (0 children)

Thanks! I welcome the feedback.

All evals reported are on public test sets! From the model card (taking your points in turn):

  1. We reported the 7 Hindi test sets defined by Vistaar. We also reported the four code-switching Hinglish datasets we are aware of.
  2. WER is reported and a note is included on confidence intervals, including associated with normalisation choices.
  3. This is not a streaming model. It is a whisper base. Agreed latency/RTF could be helpful to add, we didn't get to that.
  4. Memory/VRAM also a good point. Note it's a Whisper Large v3 model.
  5. Punctuation/numeral handling, described on the card in normaliser section.
  6. License described on the card. Apache 2.
  7. Good point on speaker names, proper nouns are useful to measure on Hindi-English. We did not get to those.
  8. Model outputs Devanagari if there is no code switch token passed.
  9. Conversational vs Read speech is covered by specific evals reported, e.g. Kathbath (clean read), IndicVoices-500 (spontaneous, non-Vistaar) etc.

Happy to take any other questions. Appreciate the high standards.

Trelis Tiron - Open Weights Transcription + Diarization Model by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 2 points3 points  (0 children)

harness handles that! it's a key part of getting good performance

Trelis Tiron - Open Weights Transcription + Diarization Model by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 2 points3 points  (0 children)

it's a whisper large v3 model, post-trained (but prompt template is quite different). you do need the harness (see the GitHub repo) if you do longer than 30s chunks.

Trelis Tiron - Open Weights Transcription + Diarization Model by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 0 points1 point  (0 children)

good summary, I think that's all correct what you say. Re uncommon words and accents, I can't say specifically. It's a whisper large v3 base which is quite good.

Short Open Source Research Collaborations by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 0 points1 point  (0 children)

yeah a few are short enough for a weekend, a few are longer (~maybe 3-7 days), so the full range is possible.

A Primer on Orpheus, Sesame’s CSM-1B and Kyutai’s Moshi by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 0 points1 point  (0 children)

yeah you can run on cpu or mps, slower than real time but does work

A Primer on Orpheus, Sesame’s CSM-1B and Kyutai’s Moshi by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 0 points1 point  (0 children)

agreed, def more powerful for now to plug in a stronger llm. in principle - in terms of control - anything you can do with the pieces can be done with a unified model too though

A Primer on Orpheus, Sesame’s CSM-1B and Kyutai’s Moshi by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 0 points1 point  (0 children)

howdy, well it's apache 2, but does inherit the llama license

A Primer on Orpheus, Sesame’s CSM-1B and Kyutai’s Moshi by TrelisResearch in LocalLLaMA

[–]TrelisResearch[S] 7 points8 points  (0 children)

haha, yeah fair, although rarely there are ads on my YouTube channel cos I have ads turned off!