Course navigation
Lesson undefined

Official resources: AssemblyAI · AssemblyAI documentation

What is AssemblyAI?

AssemblyAI is a developer-focused Voice AI platform. Instead of being primarily a consumer chat application, it provides APIs and models that developers can integrate into applications that need to transcribe, understand, and act on spoken audio.

Its platform covers pre-recorded speech-to-text, real-time transcription, Speech Understanding, Voice Agents, Dictation, Sync Speech-to-Text, guardrails, and related Voice AI infrastructure.

🎙️
Transcribe
Convert spoken audio into text.
🧠
Understand
Extract meaning and structure from conversations.
🤖
Act
Power agents, applications, analytics, and workflows.

AssemblyAI describes itself as Voice AI infrastructure for builders, with APIs for pre-recorded and real-time speech-to-text and voice-agent workflows.

What can you build with AssemblyAI?

📝
Transcription
Turn podcasts, meetings, calls, interviews, and other recordings into structured text.
⚡
Live transcription
Stream audio and receive partial and final transcripts for real-time applications.
👥
Speaker labeling
Determine who said what in multi-speaker conversations.
📊
Conversation analytics
Extract sentiment, entities, topics, key phrases, summaries, and other signals.
🤖
Voice agents
Use real-time speech infrastructure as part of production conversational agents.
🔒
Safety and privacy
Apply capabilities such as PII redaction and content moderation to speech data.

AssemblyAI workflow

1
Choose the input
Use a file, URL, short clip, or live audio stream depending on the API.
2
Select the model
Choose a speech model appropriate for accuracy, language, latency, and workload.
3
Configure context
Add key terms, prompting, language detection, speaker labels, or other options.
4
Transcribe
Send audio to the appropriate AssemblyAI endpoint or SDK.
5
Understand
Add summaries, sentiment, entities, topics, chapters, PII redaction, or other intelligence.
6
Ship the result
Store the structured output or feed it into search, analytics, an LLM, or a voice workflow.

The important architectural idea is that AssemblyAI can handle multiple speech processing steps without requiring you to assemble a separate model pipeline for every feature.

Pre-recorded Speech-to-Text

The pre-recorded Speech-to-Text API is designed for audio that already exists: podcasts, recorded meetings, interviews, calls, videos, and uploaded files.

Python SDK — basic transcription
import assemblyai as aai

aai.settings.api_key = "YOUR_API_KEY"

transcriber = aai.Transcriber()
transcript = transcriber.transcribe("meeting.mp3")

print(transcript.text)
Result
The application sends the recording to AssemblyAI and receives a transcript that can then be stored, displayed, searched, summarized, or passed to another AI system.
Conceptual SDK workflow based on AssemblyAI's documented Python examples

AssemblyAI's current pre-recorded API supports Universal-3.5 Pro and Universal-2, with features such as language detection, speaker labels, key terms, and Speech Understanding layers.

Universal-3.5 Pro

Universal-3.5 Pro is AssemblyAI's current flagship speech model for pre-recorded transcription and is also available in a dedicated real-time variant. The pre-recorded model is designed for difficult, domain-specific audio and supports natural-language prompting and precise entity handling.

🎯
Contextual control
Use natural-language context to help the model interpret domain-specific speech.
🏷️
Entity handling
Improve recognition of names, products, technical terms, and other important entities.
🌍
Multilingual
The current pre-recorded Universal-3.5 Pro documentation lists 18 languages with native code switching.
🔊
Real-world audio
The model is positioned for noisy environments, accents, and technical vocabulary.

AssemblyAI currently lists Universal-3.5 Pro as its flagship pre-recorded model, while Universal-2 remains available for high-volume multilingual transcription.

Prompting speech-to-text

A powerful AssemblyAI concept is that modern speech-to-text does not always have to be treated as a completely fixed transcription step. Universal-3.5 Pro supports natural-language prompting, allowing developers to provide context that helps the model interpret specialized audio.

DomainExplain what kind of conversation or audio this is.
EntitiesProvide context for important names, products, acronyms, or technical terms.
Expected languageDescribe the language or code-switching context when relevant.
Output needsExplain formatting or interpretation requirements when supported.
Contextual prompting
prompt = """
This is a technical support call
about PostgreSQL and Kubernetes.
Expect product names, acronyms,
and infrastructure terminology.
"""
Result
Context can help the speech model distinguish specialized vocabulary and recognize important entities more reliably than a generic transcription setup.
Illustrative prompting pattern; exact API configuration depends on the SDK/model

AssemblyAI's current prompting guide specifically documents natural-language prompting for Universal-3.5 Pro and describes it as a way to improve transcription accuracy without fine-tuning.

Speaker diarization

Speaker diarization answers a simple but important question: who said what? Instead of returning one block of text, the transcript can be divided into speaker-attributed utterances with timestamps.

Speaker-labeled transcript
Speaker A: Welcome to the meeting.
Speaker B: Thanks. I have three updates.
Speaker A: Let's start with the first one.
Result
Speaker labels make transcripts useful for meetings, interviews, call analytics, AI notetakers, compliance workflows, and downstream summarization.
Illustrative diarized transcript

AssemblyAI documents speaker diarization for pre-recorded and streaming workflows, with current feature pages describing timestamps and speaker-attributed utterances.

Speaker identification

Diarization can tell you that Speaker A and Speaker B spoke. Speaker identification goes further by using available audio context and custom profiles to associate speakers with names.

👤
Diarization
Separate the conversation into speakers.
🏷️
Identification
Associate supported speakers with known identities or profiles.
📌
Attribution
Use named speakers in summaries, analytics, and compliance workflows.

AssemblyAI's Speech Understanding API documents Speaker Identification as a purpose-built model that supports custom profiles and multi-speaker conversations.

Speech Understanding API

Transcription is only the first layer. The Speech Understanding APIturns transcripts into structured intelligence using multiple specialized models. Features can be combined in a single request.

📝
Summarization
Create concise summaries of spoken content.
😊
Sentiment
Detect sentence-level positive, neutral, or negative sentiment.
🔑
Key phrases
Extract important concepts and phrases.
🏢
Entities
Detect names, organizations, locations, dates, and other entities.
🏷️
Topics
Classify content against the IAB topic taxonomy.
📚
Auto chapters
Break long recordings into timestamped sections with summaries.
📋
Action items
Extract follow-up tasks and actions from conversations.
🌐
Translation
Translate supported speech content as part of the processing workflow.
🔒
PII redaction
Reduce exposure of sensitive personal information in transcript data.

AssemblyAI currently describes nine Speech Understanding models that can be mixed and matched on the same API request.

Real-time Speech-to-Text

For live applications, AssemblyAI provides a Realtime Speech-to-Text API. Instead of uploading a completed recording and waiting for a final result, the application streams audio through a WebSocket and receives partial and final transcript events.

Real-time streaming concept
client.connect({
  speech_model: "universal-3-5-pro",
  sample_rate: 16000,
  continuous_partials: true
});

client.stream(microphone);
Result
A live application can update its UI as speech arrives, enabling captions, notetakers, voice agents, call analytics, and interactive assistants.
Simplified representation of the documented streaming workflow

AssemblyAI currently documents approximately 150 ms P50 latency for Universal-3.5 Pro Realtime, with partial and final transcripts and support for streaming speaker diarization.

Real-time vs. pre-recorded transcription

RequirementPre-recordedRealtime
InputExisting audio/videoLive audio stream
OutputCompleted transcriptPartial + final transcript events
Typical usePodcasts, meetings, recorded callsCaptions, voice agents, live notetakers
ConnectionAPI request / job workflowSecure WebSocket stream

Product capabilities and model availability change over time; check the current AssemblyAI documentation when implementing a production integration.

Sync Speech-to-Text API

AssemblyAI also provides a Sync Speech-to-Text API for short clips. It is designed for a simpler request-response workflow where the application sends a short audio clip and receives the finished transcript without managing a polling job or WebSocket connection.

📤
Send
Submit a short supported audio clip.
⚡
Wait briefly
The API processes the clip synchronously.
📥
Receive
Read the completed transcript from the response.

AssemblyAI's current product overview describes Sync Speech-to-Text as a single request/response API for short clips, with up to two minutes per request.

Voice Agent API

AssemblyAI has expanded beyond transcription into Voice Agent APIinfrastructure. The goal is to provide the speech layer required for production conversational agents, including real-time speech recognition, turn detection, and interruption handling.

🎙️
Hear
Capture and transcribe what the user says in real time.
⏸️
Turn detection
Determine when the user has finished speaking so the agent can respond naturally.
↩️
Interruptions
Handle users speaking over the agent instead of treating every interaction as a rigid turn.

AssemblyAI lists Voice Agent API as an end-to-end voice-agent infrastructure product built around its real-time speech capabilities.

Dictation API

The Dictation API is designed for applications where speech should become polished, formatted text rather than a raw transcript. AssemblyAI describes it as combining speech-to-text with an LLM cleanup pass in one call.

🎤
Raw speech
The user speaks naturally, including conversational disfluencies.
✨
Clean output
The API returns formatted text suitable for dictation-oriented interfaces.

The current AssemblyAI product overview describes Dictation as transcription plus an LLM cleanup pass, with formatting instructions controlled by the developer.

PII redaction and content safety

Speech applications often process names, phone numbers, email addresses, account information, and other sensitive data. AssemblyAI provides safety-oriented features that can be applied as part of the speech processing pipeline.

PII Redaction
Protect sensitive transcript data

Reduce exposure of personally identifiable information in transcript output before storing or passing it downstream.

Content Moderation
Classify potentially problematic content

Add moderation-oriented analysis to speech processing workflows.

Profanity Filtering
Control transcript output

Use supported filtering capabilities when an application needs cleaner transcript text.

Compliance
Production-oriented infrastructure

AssemblyAI documents SOC 2 Type 2 and GDPR support, with BAA availability for eligible use cases.

Current platform documentation lists PII redaction and content moderation among its speech features and describes SOC 2 Type 2, GDPR, and BAA availability.

Language support and code-switching

Language requirements matter when building Voice AI. AssemblyAI's current pre-recorded Universal-2 offering covers 99 languages, while Universal-3.5 Pro focuses on high-accuracy multilingual use cases with 18 languages and native code-switching.

🌎
99 languages
Universal-2 is documented for broad multilingual transcription coverage.
🔀
Code-switching
Universal models can handle supported language switching in a conversation.
🇮🇳
Hindi
Hindi is included among the currently listed Universal-3.5 Pro Realtime languages.
🎯
Context
Language detection and prompting can help applications handle varied speech contexts.

Current AssemblyAI product pages list 99 languages for Universal-2 and 18 languages for Universal-3.5 Pro, with code-switching support described for the relevant models.

Using AssemblyAI with LLMs

AssemblyAI can serve as the speech layer in a larger AI application. A common architecture is:

1
Audio

The user speaks into a microphone or an application receives a recorded audio file.

microphonecall recordingpodcastmeeting
2
AssemblyAI

Speech-to-text converts the audio into structured language data and optional speech intelligence.

transcriptionspeakersentitiessentiment
3
LLM

An LLM can reason over the transcript, summarize it, answer questions, or decide what action should happen next.

GPTClaudeGeminiRAG
4
Application

The application displays the result, stores it, searches it, or triggers a business workflow.

CRMnotetakersupportanalytics

AssemblyAI currently also offers an LLM Gateway for calling multiple LLM providers through a unified API.

Example: AI meeting notetaker

One of the clearest applications is an AI meeting assistant. AssemblyAI can handle the speech-processing layer while an LLM or application layer turns the transcript into useful meeting artifacts.

🎙️
Record
Capture the meeting audio.
👥
Identify
Separate and identify speakers where supported.
📝
Transcribe
Generate the searchable transcript.
📌
Summarize
Extract summary, chapters, and action items.
Meeting intelligence pipeline
audio
  ↓
AssemblyAI
  ↓
speaker labels
  ↓
summary + action items
  ↓
LLM / database / UI
Result
Instead of building speech recognition, speaker separation, summarization, and extraction independently, the application can compose AssemblyAI's speech capabilities and then pass structured results to downstream systems.
Illustrative architecture for a meeting-notetaker application

Example: Voice AI agent

For a voice agent, latency becomes as important as transcription quality. The application needs to detect speech quickly, recognize when a turn has ended, send the user's intent to the agent logic, and respond without awkward delays.

🎙️ User speaks→⚡ Realtime STT→🧠 Agent / LLM→🔊 Voice response

AssemblyAI's Realtime product is positioned for voice agents and documents partial and final transcript events, turn detection, keyterm prompting, and speaker diarization.

Pricing overview

AssemblyAI uses usage-based pricing, with rates depending on the model and feature. Current published examples include:

Universal-2
$0.15 / hour

Current published pay-as-you-go rate for pre-recorded speech-to-text.

Universal-3.5 Pro
$0.21 / hour

Current published pay-as-you-go rate for pre-recorded transcription.

U3.5 Pro Realtime
$0.45 / hour

Current published pay-as-you-go rate for the flagship realtime model.

Universal Streaming
$0.15 / hour

Current published pay-as-you-go rate for the English realtime model.

These are current published rates retrieved for this lesson and may change. Optional Speech Understanding models and other features can have separate pricing. Always verify the current pricing page before implementing cost assumptions.

Important limitations

⚠️
Speech can be ambiguous
Background noise, overlapping speakers, accents, and unclear audio can still create transcription errors.
⚠️
Context matters
Specialized terminology can require key terms or contextual prompting for best results.
⚠️
Real-time is different
Streaming systems optimize for low latency, so partial transcripts should not always be treated as final text.
⚠️
AI output needs validation
Summaries, sentiment, entities, and action items are model outputs and should be reviewed for high-stakes workflows.

Best practices for developers

  • ✓ Keep API keys on the server; do not expose secrets in browser code.
  • ✓ Choose pre-recorded, realtime, or sync APIs according to the application's latency needs.
  • ✓ Add key terms or contextual prompting when the domain contains specialized vocabulary.
  • ✓ Enable speaker labels for conversations where attribution matters.
  • ✓ Treat partial realtime transcripts differently from final transcripts.
  • ✓ Add Speech Understanding features only when they provide a clear product benefit.
  • ✓ Redact or protect sensitive transcript data before storing or sharing it.
  • ✓ Monitor latency, transcription quality, usage, and cost in production.
  • ✓ Validate AI-generated summaries and extracted actions before using them for important decisions.

For developers building with audio

AssemblyAI is especially useful after students understand general AI APIs and prompting. It demonstrates how a specialized AI model can become one component in a larger production architecture: audio → speech model → structured intelligence → LLM → application.

Quick AssemblyAI checklist

  • ✓ Decide whether the audio is pre-recorded, realtime, or a short synchronous clip.
  • ✓ Choose the speech model based on accuracy, language, and latency requirements.
  • ✓ Add key terms and context for domain-specific audio.
  • ✓ Enable speaker diarization when speaker attribution matters.
  • ✓ Add Speech Understanding features such as summaries, sentiment, entities, or chapters as needed.
  • ✓ Use realtime APIs for interactive voice experiences.
  • ✓ Protect API keys and sensitive audio/transcript data.
  • ✓ Validate downstream AI outputs before using them in high-stakes workflows.