Advertisement

Empower Train Crew: DSB’s Speech-to-Text Solution – Thor Steen Larsen, Danske Statsbaner (DSB)

Session Outline

By enabling train operators and service personnel to speak directly into their devices, reports are automatically generated and classified, significantly reducing the time and complexity involved in reporting faults. The session will also cover the current state of LLMs tech in practice, including implementations using Whisper, GPT and Mistral and the steps being taken to refine this solution for better accuracy and integration.

Key Takeaways

  • Deployment of DSB’s speech-to-text solution allows train crew to report faults verbally, streamlining the process and reducing the need for manual input – this build on a robust design and infrastructure stack.
  • Current Challenges of speech to text models in Danish in noisy settings. Whisper based models face challenges with number recognition and understanding DSB-specific terminology, with an accuracy rate in Danish.
  • DSB is currently collecting more data and refining models to improve transcription accuracy, aiming for a lower Word Error Rate and a categorization accuracy of 95%.
  • Finetuning open-source models DSB are getting higher accuracy on transcription and want to give back to the open-source community.
Add a comment

Leave a Reply