Skip to main content
The ElevenLabs Speech to Text node transcribes audio into text using the ElevenLabs API. It supports automatic language detection, speaker diarization (identifying different speakers), and audio event tagging (annotating sounds like laughter or music in the transcript).

Inputs

Common Inputs

Scribe v2 Inputs

These parameters are shown when the "scribe_v2" model is selected. Note: num_speakers cannot be set to a value greater than 0 when diarize is enabled. You must either disable diarize or set num_speakers to 0.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 7eb5d72615aa8a9e4a8014e45b39cf83dc8d8432d7ce0dccba20489be80a5830