top of page

FREQUENTLY ASKED QUESTIONS

Clear answers about AI audio description, accessibility compliance, and how Audible Sight works — for enterprise, education, and government teams.

Audio description basics

What is audio description?

A narration track that describes what’s on screen for viewers who are blind or have low vision.

It covers actions, settings, on-screen text, graphics and facial expressions. It never talks over dialogue — descriptions play in natural pauses, or in pauses inserted to fit them.

What is AI audio description?

The same narration, generated automatically instead of by a human writer, voice actor and video editor.

Computer vision drafts the descriptions and speech synthesis voices them. You review and approve before anything is published.

What’s the difference between standard, extended and hybrid description?

Standard fits into existing pauses. Extended pauses the video to make room. Hybrid uses both.

  • Standard: descriptions fit into natural gaps in dialogue. Typical use: movies, TV.

  • Extended: the video briefly pauses when gaps are too short. Typical use: education, training, informational video.

  • Hybrid: standard where it fits, extended where it doesn’t. Typical use: mixed content.

Audible Sight picks the best mode for each video automatically.

How is audio description different from captions?

Captions turn sound into text. Audio description turns visuals into sound.

They serve different audiences (deaf/hard of hearing vs. blind/low vision) and are separate legal requirements — providing one does not satisfy the other. Audible Sight produces both.

Does every video need audio description?

Most do.

If someone can follow the video without seeing the screen and miss nothing, it may not need description. In practice, instructional video, slide presentations, promotional video, and anything with on-screen text almost always does.

Should we use a tool to decide which videos need description first?

We don’t recommend it — describing everything is one step; triage makes it two.

  • Every video already goes through a tool for captions. Audible Sight does captions and description in the same pass.

  • Triage tools output a recommendation, not an accessible video — you still have to describe the list.

  • The risks aren’t equal: describing a video that didn’t need it costs almost nothing; skipping one that did leaves an inaccessible video and a record that you chose not to describe it.

  • Videos that need little description simply get very little. The sorting happens as a byproduct.

Do descriptions have to line up exactly with what they describe?

No. Close is fine — and often better.

If dialogue blocks the exact moment, the description moves to a nearby gap. Combining two descriptions into one gap also means fewer interruptions. Audible Sight handles placement automatically, and you can move anything by drag and drop.

How do viewers access the described version?

Two common options.

  1. Separate version: post it next to the original with “(audio described)” added to the title, as a link or thumbnail.

  2. Alternate audio track: works only if there are no extended descriptions and your player supports alternate audio tracks.

What mistakes do beginners make?

Over-describing, and over-syncing.

  • Too many descriptions, too much detail, flowery language. Less is more.

  • Trying to sync every description exactly to its image. Grouping descriptions into one gap is often preferred.

Compliance & legal

Is audio description legally required?

Yes, for many organizations.

  • US: Section 508, Section 504 (federally funded entities), the ADA, and the CVAA.

  • EU: the European Accessibility Act, in force since June 28, 2025.

What does ADA Title II require, and when?

WCAG 2.1 AA for state and local government web content — deadlines in 2027 and 2028.

The DOJ’s 2024 rule covers public colleges and universities, K-12 districts, cities, towns and special districts. For prerecorded video, that means captions (1.2.2) and audio description (1.2.5).

  • Population 50,000+: April 26, 2027

  • Smaller entities and special districts: April 26, 2028

DOJ extended the original dates by one year in April 2026.

Does this apply to old videos?

Yes, if they’re still in use.

Only genuinely archived video that’s no longer used or posted is exempt. Anything live on your website, LMS, social channels or media library is in scope, whenever it was made.

What about organizations that aren’t government?

Other laws and contracts apply.

  • Section 508: federal agencies and their contractors

  • Section 504: organizations receiving federal funding

  • ADA Title III: places of public accommodation

  • Private companies, publishers, media: contracts, distribution partners and their own accessibility commitments

Does Audible Sight meet WCAG, Section 508 and ADA requirements for video?

Yes.

  • WCAG 2.1 AA: captions (1.2.2) + audio description (1.2.5) — covered

  • ADA Title II: WCAG 2.1 AA — covered

  • Section 508 (2018 refresh): WCAG 2.0 AA, same video criteria — covered

  • CVAA, European Accessibility Act: audio description — covered

A VPAT is available for procurement. Full-site conformance depends on more than video, but for video, this is what the standards ask for.

Which WCAG criteria does Audible Sight cover?

All four prerecorded-video criteria, A through AAA.

  • Captions: WCAG 1.2.2, Level A

  • Standard audio description: WCAG 1.2.5, Level AA

  • Extended audio description: WCAG 1.2.7, Level AAA

  • Descriptive transcript: WCAG 1.2.8, Level AAA

These are the video requirements referenced by ADA Title II, Sections 508 and 504, the EAA, EN 301 549, and Canadian accessibility law.

Is AI-generated description acceptable for compliance?

Yes. WCAG sets a standard for the output, not how it’s made.

Your team reviews and approves every description before export, so a person is accountable for what’s published — the same model already accepted for auto-generated captions with human review.

Is there an accuracy percentage we need to hit?

No. WCAG requires description to be present, accurate and synchronized — no percentage.

See “How accurate is Audible Sight?” below.

How it works

What does the platform do?

Drafts descriptions and captions, lets you edit them, then exports a finished video.

  1. Upload a video. Audible Sight splits it into scenes and drafts descriptions and captions within minutes.

  2. Review and edit the text; reposition descriptions by drag and drop.

  3. Preview, then export. Descriptions are voiced and inserted at the right points.

No description writer, voice actor or video editor needed.

Is it software or a service?

Software (SaaS). Nothing to install.

You upload, review and export on your own schedule — no vendor queue, no waiting days or weeks, no change-request rounds. A backlog of hundreds of hours becomes something you can schedule against a deadline.

Do we need accessibility, video or audio expertise?

No.

Descriptions are generated for you, and Audible Sight places them in the video. No writing, voice recording or editing skills required.

How long does it take?

Drafts are ready within minutes.

Traditional description takes a human 30–60 minutes of work per 5 minutes of video, plus recording and editing — days or weeks end to end.

What file types go in, and what comes out?

MP4 or MOV in. MP4, audio track, captions and transcript out.

  • Video with description mixed in: .mp4 (4K or 1080p)

  • Description audio only: separate audio file

  • Captions: .srt or .vtt

  • Descriptive transcript (WCAG 1.2.8): .txt

All outputs are included at no extra charge. Use the video anywhere the original goes — YouTube, Vimeo, Kaltura, Panopto, Canvas, Brightspace, web pages, social, PowerPoint.

Descriptive transcripts also help deafblind users on braille displays and make video searchable.

Are there limits on file size, length, resolution or monthly use?

No.

No caps on file size, length, resolution (4K included) or monthly usage. Your full credit allotment is available from day one, so you can clear a backlog in your first month — useful ahead of the April 2027 Title II deadline (extended from 2026), when remediation work tends to arrive in bulk.

Can we upload in batches?

Yes.

  • Enterprise license: unlimited concurrent uploads

  • Pro license: 2 concurrent uploads

How many user accounts can we have? Can they share?

Enterprise licenses include unlimited accounts. Pro licenses include up to 6.

Staff can share files and credits.

What does it integrate with?

Single sign-on, Kaltura and Google Drive.

Custom integrations are added as needed.

Do you do live description for events or Zoom?

No — recorded video only.

WCAG’s description criteria apply to prerecorded media, so the DOJ web rule doesn’t require live description (though ADA Title II has broader effective-communication duties for live events).

If you post recordings of live sessions, those are in scope — and that’s exactly what Audible Sight handles.

Can it handle film and TV?

Yes, with Audible Sight Studio Edition.

Built for precise placement in films, episodes and other entertainment. Contact us for details.

Editing & quality

How accurate is Audible Sight?

Accurate on what can be checked; you decide the rest.

Captions can claim 99% accuracy because there’s one right answer. Description isn’t like that — two skilled describers produce two different, correct descriptions.

What we hold ourselves to:

  • Describes what’s actually on screen

  • Captures on-screen text that isn’t spoken

  • Fits gaps without talking over dialogue

  • Sits close to what it describes

What’s your call: how much detail, and what to prioritize. That’s why you get an editable draft, not a locked file.

Why not fully automatic, with no human review?

Because AI still makes mistakes, and unreviewed description fails silently.

A bad caption is usually obvious. A description that confidently narrates something not on screen reads perfectly fluently — only the person who depends on it notices. Keeping review with your team is the design, not a limitation.

What can we change in the drafts?

Everything.

  • Edit wording in a simple text editor

  • Choose which scenes to describe; deselect redundant descriptions

  • Add a description anywhere with Add additional description

  • Drag to reposition or group descriptions on a timeline

  • Preview before exporting

How do we fix mispronounced names or terms?

Two ways.

  • One-off: respell the word phonetically in the description text.

  • Recurring terms: send the word and pronunciation to wehearyou@audiblesight.ai. Our I Now Pronounce You feature adds it, usually within one business day, for every future video.

Voices & languages

How many voices and languages are there?

14 languages on every license. Pro licenses include 50 voices; Enterprise licenses include all 136.

  • English — US, UK, Canada, Australia, South Africa, Mexico, Ireland, India

  • Spanish — US, Mexico, Spain

  • French — France, Canada, Belgium

  • Portuguese — Brazil, Portugal

  • German, Italian, Japanese, Korean, Mandarin, Cantonese

Male and female voices in a range of styles and accents, with two speaking speeds. More are added periodically.

Where do the voices come from?

Licensed from real, paid voice actors.

All voices come through WellSaid Labs, built from consenting, compensated performers who earn royalties on every use — not scraped recordings.

Pricing & getting started

How much does it cost?

A Pro license is $99/month, billed annually ($1,188/year), and includes 1,000 minutes.

  • Pro license: $1,188/year · 1,000 minutes · about $1.19/minute

  • Enterprise license: higher volume, lower per-minute cost — contact us for a quote

  • Extra minutes: $1/minute, or renew any time

Customers report saving 80–90%+ of the cost and time of traditional description.

Are there discounts?

Yes — for education, non-profit, government and volume.

These apply to Enterprise licenses. Contact us for details.

Do unused minutes expire?

No.

They’re Forever Minutes — they roll over for as long as you’re subscribed.

Can we try it first?

Yes — a free trial with 10 minutes of video.

How do we get started?

Start a free trial, or contact us for an Enterprise license.

  1. Create a trial account at audiblesight.ai. Upgrade to a Pro license any time from the Account link.

  2. For an Enterprise license, email info@audiblesight.ai.

Who uses Audible Sight?

Organizations that own video — not individual consumers.

Government, education, businesses, publishers, media and non-profits, in the US and abroad.

Support, training & security

What support is included?

Pro licenses include support weekdays 8am–7pm ET. Enterprise licenses include 24/7 premier support and a dedicated success manager.

All support is US-based.

Is training required?

No.

Most customers start without training. Optional resources:

  • Short tutorials on the Audible Sight YouTube channel

  • 2 hours of self-study content

  • 45-minute live onboarding (Enterprise licenses)

Is our video used to train AI?

No.

Customer videos and generated descriptions are never used to train Audible Sight’s models.

Where is our data stored?

Microsoft Azure, East US.

  • AES-256 encryption at rest

  • TLS 1.2+ in transit

  • Geo-redundant continuous backup

Still have a question?

Who do we contact?

bottom of page