Official resources: AssemblyAI · AssemblyAI documentation
What is AssemblyAI?
AssemblyAI is a developer-focused Voice AI platform. Instead of being primarily a consumer chat application, it provides APIs and models that developers can integrate into applications that need to transcribe, understand, and act on spoken audio.
Its platform covers pre-recorded speech-to-text, real-time transcription, Speech Understanding, Voice Agents, Dictation, Sync Speech-to-Text, guardrails, and related Voice AI infrastructure.
AssemblyAI describes itself as Voice AI infrastructure for builders, with APIs for pre-recorded and real-time speech-to-text and voice-agent workflows.
What can you build with AssemblyAI?
AssemblyAI workflow
The important architectural idea is that AssemblyAI can handle multiple speech processing steps without requiring you to assemble a separate model pipeline for every feature.
Pre-recorded Speech-to-Text
The pre-recorded Speech-to-Text API is designed for audio that already exists: podcasts, recorded meetings, interviews, calls, videos, and uploaded files.
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
transcriber = aai.Transcriber()
transcript = transcriber.transcribe("meeting.mp3")
print(transcript.text)AssemblyAI's current pre-recorded API supports Universal-3.5 Pro and Universal-2, with features such as language detection, speaker labels, key terms, and Speech Understanding layers.
Universal-3.5 Pro
Universal-3.5 Pro is AssemblyAI's current flagship speech model for pre-recorded transcription and is also available in a dedicated real-time variant. The pre-recorded model is designed for difficult, domain-specific audio and supports natural-language prompting and precise entity handling.
AssemblyAI currently lists Universal-3.5 Pro as its flagship pre-recorded model, while Universal-2 remains available for high-volume multilingual transcription.
Prompting speech-to-text
A powerful AssemblyAI concept is that modern speech-to-text does not always have to be treated as a completely fixed transcription step. Universal-3.5 Pro supports natural-language prompting, allowing developers to provide context that helps the model interpret specialized audio.
prompt = """
This is a technical support call
about PostgreSQL and Kubernetes.
Expect product names, acronyms,
and infrastructure terminology.
"""AssemblyAI's current prompting guide specifically documents natural-language prompting for Universal-3.5 Pro and describes it as a way to improve transcription accuracy without fine-tuning.
Speaker diarization
Speaker diarization answers a simple but important question: who said what? Instead of returning one block of text, the transcript can be divided into speaker-attributed utterances with timestamps.
Speaker A: Welcome to the meeting.
Speaker B: Thanks. I have three updates.
Speaker A: Let's start with the first one.AssemblyAI documents speaker diarization for pre-recorded and streaming workflows, with current feature pages describing timestamps and speaker-attributed utterances.
Speaker identification
Diarization can tell you that Speaker A and Speaker B spoke. Speaker identification goes further by using available audio context and custom profiles to associate speakers with names.
AssemblyAI's Speech Understanding API documents Speaker Identification as a purpose-built model that supports custom profiles and multi-speaker conversations.
Speech Understanding API
Transcription is only the first layer. The Speech Understanding APIturns transcripts into structured intelligence using multiple specialized models. Features can be combined in a single request.
AssemblyAI currently describes nine Speech Understanding models that can be mixed and matched on the same API request.
Real-time Speech-to-Text
For live applications, AssemblyAI provides a Realtime Speech-to-Text API. Instead of uploading a completed recording and waiting for a final result, the application streams audio through a WebSocket and receives partial and final transcript events.
client.connect({
speech_model: "universal-3-5-pro",
sample_rate: 16000,
continuous_partials: true
});
client.stream(microphone);AssemblyAI currently documents approximately 150 ms P50 latency for Universal-3.5 Pro Realtime, with partial and final transcripts and support for streaming speaker diarization.
Real-time vs. pre-recorded transcription
| Requirement | Pre-recorded | Realtime |
|---|---|---|
| Input | Existing audio/video | Live audio stream |
| Output | Completed transcript | Partial + final transcript events |
| Typical use | Podcasts, meetings, recorded calls | Captions, voice agents, live notetakers |
| Connection | API request / job workflow | Secure WebSocket stream |
Product capabilities and model availability change over time; check the current AssemblyAI documentation when implementing a production integration.
Sync Speech-to-Text API
AssemblyAI also provides a Sync Speech-to-Text API for short clips. It is designed for a simpler request-response workflow where the application sends a short audio clip and receives the finished transcript without managing a polling job or WebSocket connection.
AssemblyAI's current product overview describes Sync Speech-to-Text as a single request/response API for short clips, with up to two minutes per request.
Voice Agent API
AssemblyAI has expanded beyond transcription into Voice Agent APIinfrastructure. The goal is to provide the speech layer required for production conversational agents, including real-time speech recognition, turn detection, and interruption handling.
AssemblyAI lists Voice Agent API as an end-to-end voice-agent infrastructure product built around its real-time speech capabilities.
Dictation API
The Dictation API is designed for applications where speech should become polished, formatted text rather than a raw transcript. AssemblyAI describes it as combining speech-to-text with an LLM cleanup pass in one call.
The current AssemblyAI product overview describes Dictation as transcription plus an LLM cleanup pass, with formatting instructions controlled by the developer.
PII redaction and content safety
Speech applications often process names, phone numbers, email addresses, account information, and other sensitive data. AssemblyAI provides safety-oriented features that can be applied as part of the speech processing pipeline.
Reduce exposure of personally identifiable information in transcript output before storing or passing it downstream.
Add moderation-oriented analysis to speech processing workflows.
Use supported filtering capabilities when an application needs cleaner transcript text.
AssemblyAI documents SOC 2 Type 2 and GDPR support, with BAA availability for eligible use cases.
Current platform documentation lists PII redaction and content moderation among its speech features and describes SOC 2 Type 2, GDPR, and BAA availability.
Language support and code-switching
Language requirements matter when building Voice AI. AssemblyAI's current pre-recorded Universal-2 offering covers 99 languages, while Universal-3.5 Pro focuses on high-accuracy multilingual use cases with 18 languages and native code-switching.
Current AssemblyAI product pages list 99 languages for Universal-2 and 18 languages for Universal-3.5 Pro, with code-switching support described for the relevant models.
Using AssemblyAI with LLMs
AssemblyAI can serve as the speech layer in a larger AI application. A common architecture is:
The user speaks into a microphone or an application receives a recorded audio file.
Speech-to-text converts the audio into structured language data and optional speech intelligence.
An LLM can reason over the transcript, summarize it, answer questions, or decide what action should happen next.
The application displays the result, stores it, searches it, or triggers a business workflow.
AssemblyAI currently also offers an LLM Gateway for calling multiple LLM providers through a unified API.
Example: AI meeting notetaker
One of the clearest applications is an AI meeting assistant. AssemblyAI can handle the speech-processing layer while an LLM or application layer turns the transcript into useful meeting artifacts.
audio
↓
AssemblyAI
↓
speaker labels
↓
summary + action items
↓
LLM / database / UIExample: Voice AI agent
For a voice agent, latency becomes as important as transcription quality. The application needs to detect speech quickly, recognize when a turn has ended, send the user's intent to the agent logic, and respond without awkward delays.
AssemblyAI's Realtime product is positioned for voice agents and documents partial and final transcript events, turn detection, keyterm prompting, and speaker diarization.
Pricing overview
AssemblyAI uses usage-based pricing, with rates depending on the model and feature. Current published examples include:
Current published pay-as-you-go rate for pre-recorded speech-to-text.
Current published pay-as-you-go rate for pre-recorded transcription.
Current published pay-as-you-go rate for the flagship realtime model.
Current published pay-as-you-go rate for the English realtime model.
These are current published rates retrieved for this lesson and may change. Optional Speech Understanding models and other features can have separate pricing. Always verify the current pricing page before implementing cost assumptions.
Important limitations
Best practices for developers
- ✓ Keep API keys on the server; do not expose secrets in browser code.
- ✓ Choose pre-recorded, realtime, or sync APIs according to the application's latency needs.
- ✓ Add key terms or contextual prompting when the domain contains specialized vocabulary.
- ✓ Enable speaker labels for conversations where attribution matters.
- ✓ Treat partial realtime transcripts differently from final transcripts.
- ✓ Add Speech Understanding features only when they provide a clear product benefit.
- ✓ Redact or protect sensitive transcript data before storing or sharing it.
- ✓ Monitor latency, transcription quality, usage, and cost in production.
- ✓ Validate AI-generated summaries and extracted actions before using them for important decisions.
For developers building with audio
AssemblyAI is especially useful after students understand general AI APIs and prompting. It demonstrates how a specialized AI model can become one component in a larger production architecture: audio → speech model → structured intelligence → LLM → application.
Quick AssemblyAI checklist
- ✓ Decide whether the audio is pre-recorded, realtime, or a short synchronous clip.
- ✓ Choose the speech model based on accuracy, language, and latency requirements.
- ✓ Add key terms and context for domain-specific audio.
- ✓ Enable speaker diarization when speaker attribution matters.
- ✓ Add Speech Understanding features such as summaries, sentiment, entities, or chapters as needed.
- ✓ Use realtime APIs for interactive voice experiences.
- ✓ Protect API keys and sensitive audio/transcript data.
- ✓ Validate downstream AI outputs before using them in high-stakes workflows.