ElevenLabs: Speech to text

This ComfyUI workflow leverages the powerful capabilities of the ElevenLabs model to transcribe speech from audio or video files into text. The workflow begins with the 'LoadAudio' or 'RecordAudio' node, which allows users to either upload an existing audio file or record new audio directly within the interface. Once the audio input is secured, the 'ElevenLabsSpeechToText' node processes the audio data, utilizing advanced speech recognition algorithms to convert spoken words into accurate text. This text is then made available for preview and editing through the 'PreviewAny' node, ensuring that users can review and refine the transcription as needed. This workflow is particularly useful for creating transcripts from interviews, lectures, or any spoken content, making it an invaluable tool for content creators, researchers, and professionals who need to convert speech to text efficiently.

API

Use this workflow from code

Every Comfy workflow is a JSON graph. The payload below is this workflow, exactly as ComfyUI runs it — fetch it from the URL, keep it in version control, or load it in ComfyUI and run it node by node.

api_elevenLabs_speech_to_text.json
Fetching workflow JSON…

Run it from TypeScript or Python with the Comfy SDK. The same code targets Comfy Cloud or a ComfyUI you host yourself — only the base URL changes.

// Install (beta)
npm i @comfyorg/sdk

// Run this workflow (TypeScript)
import { Comfy } from "@comfyorg/sdk";

const client = new Comfy({ apiKey: "comfyui-..." });
const wf = await client.workflows.fromFile("workflow_api.json");
const job = await client.run(wf);
await job.getOutputs("<output-node-id>")[0].toFile("output.png");

The SDK takes a workflow in API format: open this workflow in ComfyUI and use File → Export Workflow (API).

Comfy Cloud API access requires a plan with an API key.

FAQ

Frequently Asked Questions

View all workflows
Showing 30 of 30 templates