Audio-To-Text Conversion with Speaker Identification

Author
Keywords
Abstract

The ability to search for specific pieces of information presented by a specific person in audio meetings is a useful capability. The search becomes easy if the transcript of the meeting is annotated with speaker id. This work proposes an effective Audio-To-Text Conversion system with speaker identification. The proposed method uses two parallel pipelines where one pipeline identifies the speaker of each of the audio segments in a multi-speaker meeting session and the other pipeline converts the same audio segment in to corresponding text. The system uses a foundation model for text transcription with timestamping and uses SpeechBrain ECAPA-TDNN encoder for generating speaker embeddings that are 512-dimensional and speaker-robust. The speaker embeddings are classified using a multi-layer perceptron (MLP) classifier. By combining transcription, feature extraction, and classification, it ensures correct segmentation of dialogues and marks them as speaker labels with a confidence threshold for an "Unknown"label. The system classifies known speaker accurately with an F1 score ranging from 82% to 88% in a multi-speaker setup. It significantly enhances the readability and organization in transcribed multi-party conversations.

Year of Conference
2026
Conference Name
Proceedings of the IEEE International Conference on AI Engineering and Innovations, AIEI 2026
Publisher
Institute of Electrical and Electronics Engineers Inc.
ISBN Number
979-833156045-4 (ISBN)
URL
https://ieeexplore.ieee.org/document/11497972
DOI
10.1109/AIEI69164.2026.11497972
Short Title
Proc. IEEE Int. Conf. AI Eng. Innov., AIEI
Conference Proceedings
Download citation
Cits
0
CIT

For admissions and all other information, please visit the official website of

Cambridge Institute of Technology

Cambridge Group of Institutions

Contact

Web portal developed and administered by Dr. Subrahmanya S. Katte, Dean - Academics.

Contact the Site Admin.