FREQUENTLY ASKED QUESTIONS
Clear answers about AI audio description, accessibility compliance, and how Audible Sight works — for enterprise, education, and government teams.
Audio description basics
What is audio description?
A narration track that describes what’s on screen for viewers who are blind or have low vision.
It covers actions, settings, on-screen text, graphics and facial expressions. It never talks over dialogue — descriptions play in natural pauses, or in pauses inserted to fit them.
What is AI audio description?
The same narration, generated automatically instead of by a human writer, voice actor and video editor.
Computer vision drafts the descriptions and speech synthesis voices them. You review and approve before anything is published.
What’s the difference between standard, extended and hybrid description?
Standard fits into existing pauses. Extended pauses the video to make room. Hybrid uses both.
Standard: descriptions fit into natural gaps in dialogue. Typical use: movies, TV.
Extended: the video briefly pauses when gaps are too short. Typical use: education, training, informational video.
Hybrid: standard where it fits, extended where it doesn’t. Typical use: mixed content.
Audible Sight picks the best mode for each video automatically.
How is audio description different from captions?
Captions turn sound into text. Audio description turns visuals into sound.
They serve different audiences (deaf/hard of hearing vs. blind/low vision) and are separate legal requirements — providing one does not satisfy the other. Audible Sight produces both.
Does every video need audio description?
Most do.
If someone can follow the video without seeing the screen and miss nothing, it may not need description. In practice, instructional video, slide presentations, promotional video, and anything with on-screen text almost always does.
Should we use a tool to decide which videos need description first?
We don’t recommend it — describing everything is one step; triage makes it two.
Every video already goes through a tool for captions. Audible Sight does captions and description in the same pass.
Triage tools output a recommendation, not an accessible video — you still have to describe the list.
The risks aren’t equal: describing a video that didn’t need it costs almost nothing; skipping one that did leaves an inaccessible video and a record that you chose not to describe it.
Videos that need little description simply get very little. The sorting happens as a byproduct.
Do descriptions have to line up exactly with what they describe?
No. Close is fine — and often better.
If dialogue blocks the exact moment, the description moves to a nearby gap. Combining two descriptions into one gap also means fewer interruptions. Audible Sight handles placement automatically, and you can move anything by drag and drop.
How do viewers access the described version?
Two common options.
Separate version: post it next to the original with “(audio described)” added to the title, as a link or thumbnail.
Alternate audio track: works only if there are no extended descriptions and your player supports alternate audio tracks.
What mistakes do beginners make?
Over-describing, and over-syncing.
Too many descriptions, too much detail, flowery language. Less is more.
Trying to sync every description exactly to its image. Grouping descriptions into one gap is often preferred.
Compliance & legal
Is audio description legally required?
Yes, for many organizations.
US: Section 508, Section 504 (federally funded entities), the ADA, and the CVAA.
EU: the European Accessibility Act, in force since June 28, 2025.
What does ADA Title II require, and when?
WCAG 2.1 AA for state and local government web content — deadlines in 2027 and 2028.
The DOJ’s 2024 rule covers public colleges and universities, K-12 districts, cities, towns and special districts. For prerecorded video, that means captions (1.2.2) and audio description (1.2.5).
Population 50,000+: April 26, 2027
Smaller entities and special districts: April 26, 2028
DOJ extended the original dates by one year in April 2026.
Does this apply to old videos?
Yes, if they’re still in use.
Only genuinely archived video that’s no longer used or posted is exempt. Anything live on your website, LMS, social channels or media library is in scope, whenever it was made.
What about organizations that aren’t government?
Other laws and contracts apply.
Section 508: federal agencies and their contractors
Section 504: organizations receiving federal funding
ADA Title III: places of public accommodation
Private companies, publishers, media: contracts, distribution partners and their own accessibility commitments
Does Audible Sight meet WCAG, Section 508 and ADA requirements for video?
Yes.
WCAG 2.1 AA: captions (1.2.2) + audio description (1.2.5) — covered
ADA Title II: WCAG 2.1 AA — covered
Section 508 (2018 refresh): WCAG 2.0 AA, same video criteria — covered
CVAA, European Accessibility Act: audio description — covered
A VPAT is available for procurement. Full-site conformance depends on more than video, but for video, this is what the standards ask for.
Which WCAG criteria does Audible Sight cover?
All four prerecorded-video criteria, A through AAA.
Captions: WCAG 1.2.2, Level A
Standard audio description: WCAG 1.2.5, Level AA
Extended audio description: WCAG 1.2.7, Level AAA
Descriptive transcript: WCAG 1.2.8, Level AAA
These are the video requirements referenced by ADA Title II, Sections 508 and 504, the EAA, EN 301 549, and Canadian accessibility law.
Is AI-generated description acceptable for compliance?
Yes. WCAG sets a standard for the output, not how it’s made.
Your team reviews and approves every description before export, so a person is accountable for what’s published — the same model already accepted for auto-generated captions with human review.
Is there an accuracy percentage we need to hit?
No. WCAG requires description to be present, accurate and synchronized — no percentage.
See “How accurate is Audible Sight?” below.
How it works
What does the platform do?
Drafts descriptions and captions, lets you edit them, then exports a finished video.
Upload a video. Audible Sight splits it into scenes and drafts descriptions and captions within minutes.
Review and edit the text; reposition descriptions by drag and drop.
Preview, then export. Descriptions are voiced and inserted at the right points.
No description writer, voice actor or video editor needed.
Is it software or a service?
Software (SaaS). Nothing to install.
You upload, review and export on your own schedule — no vendor queue, no waiting days or weeks, no change-request rounds. A backlog of hundreds of hours becomes something you can schedule against a deadline.
Do we need accessibility, video or audio expertise?
No.
Descriptions are generated for you, and Audible Sight places them in the video. No writing, voice recording or editing skills required.
How long does it take?
Drafts are ready within minutes.
Traditional description takes a human 30–60 minutes of work per 5 minutes of video, plus recording and editing — days or weeks end to end.
What file types go in, and what comes out?
MP4 or MOV in. MP4, audio track, captions and transcript out.
Video with description mixed in: .mp4 (4K or 1080p)
Description audio only: separate audio file
Captions: .srt or .vtt
Descriptive transcript (WCAG 1.2.8): .txt
All outputs are included at no extra charge. Use the video anywhere the original goes — YouTube, Vimeo, Kaltura, Panopto, Canvas, Brightspace, web pages, social, PowerPoint.
Descriptive transcripts also help deafblind users on braille displays and make video searchable.
Are there limits on file size, length, resolution or monthly use?
No.
No caps on file size, length, resolution (4K included) or monthly usage. Your full credit allotment is available from day one, so you can clear a backlog in your first month — useful ahead of the April 2027 Title II deadline (extended from 2026), when remediation work tends to arrive in bulk.
Can we upload in batches?
Yes.
Enterprise license: unlimited concurrent uploads
Pro license: 2 concurrent uploads
How many user accounts can we have? Can they share?
Enterprise licenses include unlimited accounts. Pro licenses include up to 6.
Staff can share files and credits.
What does it integrate with?
Single sign-on, Kaltura and Google Drive.
Custom integrations are added as needed.
Do you do live description for events or Zoom?
No — recorded video only.
WCAG’s description criteria apply to prerecorded media, so the DOJ web rule doesn’t require live description (though ADA Title II has broader effective-communication duties for live events).
If you post recordings of live sessions, those are in scope — and that’s exactly what Audible Sight handles.
Can it handle film and TV?
Yes, with Audible Sight Studio Edition.
Built for precise placement in films, episodes and other entertainment. Contact us for details.
Editing & quality
How accurate is Audible Sight?
Accurate on what can be checked; you decide the rest.
Captions can claim 99% accuracy because there’s one right answer. Description isn’t like that — two skilled describers produce two different, correct descriptions.
What we hold ourselves to:
Describes what’s actually on screen
Captures on-screen text that isn’t spoken
Fits gaps without talking over dialogue
Sits close to what it describes
What’s your call: how much detail, and what to prioritize. That’s why you get an editable draft, not a locked file.
Why not fully automatic, with no human review?
Because AI still makes mistakes, and unreviewed description fails silently.
A bad caption is usually obvious. A description that confidently narrates something not on screen reads perfectly fluently — only the person who depends on it notices. Keeping review with your team is the design, not a limitation.
What can we change in the drafts?
Everything.
Edit wording in a simple text editor
Choose which scenes to describe; deselect redundant descriptions
Add a description anywhere with Add additional description
Drag to reposition or group descriptions on a timeline
Preview before exporting
How do we fix mispronounced names or terms?
Two ways.
One-off: respell the word phonetically in the description text.
Recurring terms: send the word and pronunciation to wehearyou@audiblesight.ai. Our I Now Pronounce You feature adds it, usually within one business day, for every future video.
Voices & languages
How many voices and languages are there?
14 languages on every license. Pro licenses include 50 voices; Enterprise licenses include all 136.
English — US, UK, Canada, Australia, South Africa, Mexico, Ireland, India
Spanish — US, Mexico, Spain
French — France, Canada, Belgium
Portuguese — Brazil, Portugal
German, Italian, Japanese, Korean, Mandarin, Cantonese
Male and female voices in a range of styles and accents, with two speaking speeds. More are added periodically.
Where do the voices come from?
Licensed from real, paid voice actors.
All voices come through WellSaid Labs, built from consenting, compensated performers who earn royalties on every use — not scraped recordings.
Pricing & getting started
How much does it cost?
A Pro license is $99/month, billed annually ($1,188/year), and includes 1,000 minutes.
Pro license: $1,188/year · 1,000 minutes · about $1.19/minute
Enterprise license: higher volume, lower per-minute cost — contact us for a quote
Extra minutes: $1/minute, or renew any time
Customers report saving 80–90%+ of the cost and time of traditional description.
Are there discounts?
Yes — for education, non-profit, government and volume.
These apply to Enterprise licenses. Contact us for details.
Do unused minutes expire?
No.
They’re Forever Minutes — they roll over for as long as you’re subscribed.
Can we try it first?
Yes — a free trial with 10 minutes of video.
How do we get started?
Start a free trial, or contact us for an Enterprise license.
Create a trial account at audiblesight.ai. Upgrade to a Pro license any time from the Account link.
For an Enterprise license, email info@audiblesight.ai.
Who uses Audible Sight?
Organizations that own video — not individual consumers.
Government, education, businesses, publishers, media and non-profits, in the US and abroad.
Support, training & security
What support is included?
Pro licenses include support weekdays 8am–7pm ET. Enterprise licenses include 24/7 premier support and a dedicated success manager.
All support is US-based.
Is training required?
No.
Most customers start without training. Optional resources:
Short tutorials on the Audible Sight YouTube channel
2 hours of self-study content
45-minute live onboarding (Enterprise licenses)
Is our video used to train AI?
No.
Customer videos and generated descriptions are never used to train Audible Sight’s models.
Where is our data stored?
Microsoft Azure, East US.
AES-256 encryption at rest
TLS 1.2+ in transit
Geo-redundant continuous backup
Still have a question?
Who do we contact?
Sales and Enterprise licenses: info@audiblesight.ai
Pronunciations & feedback: wehearyou@audiblesight.ai