Evaluation of AI-generated audio description for factual TV/media genres

Start date

January 2025

End date

Ongoing

About the project

This project aims to evaluate emerging AI solutions, including AI models and AI-based pipelines, to assess their effectiveness in generating simple audio descriptions for a selected factual TV genre. The project will explore the feasibility of leveraging generative AI to make TV content more accessible for blind and partially sighted audiences.

The demand for accessible TV content is growing, driven by increased awareness of diversity among stakeholders and expanding accessibility legislation. However, current accessibility-related practices such as audio description (AD) are costly and time-intensive, as they rely heavily on specialised human expertise. This dependency has led to significant gaps in AD provision, particularly in live TV broadcasts. At present, only around 10% of TV content is audio described, compared to over 90% of TV content that includes subtitles, while both accessibility practices could benefit diverse audiences. 

The current gaps, coupled with legislative requirements, have sparked strong interest in AI-driven AD solutions within the media and broadcast sectors. Rapid advances in generative AI hold considerable potential for automating the generation of AD, but current AI technology falls short in key areas. AI-generated AD often lacks the nuance and sophistication of human descriptions, particularly in terms of accuracy and specificity of character, object and action description, in selecting relevant information, and in terms of creating coherent, meaningful narratives across video scenes. While computer vision communities have intensified research into the automatic generation of video descriptions (‘video captioning’), little attention has been paid to the specifics of AD, such as the need to integrate AD with existing film dialogue or narration, and to maintain cohesion between AD fragments and with the overall context of a programme.

Further challenges associated with AI-generated AD include potential bias in training data, which may impact accuracy and user experience, and a lack of transparency in description generation (the black box problem in generative AI), which make it difficult to predict how AI descriptions are generated, potentially resulting in critical errors.

While this means that AI-generated AD is currently unlikely to be suitable for use without human supervision and review, feasibility testing using different AI models and pipelines is crucial at this stage. It will help identify key challenges (e.g. the areas of critical error) and opportunities for safe and systematic development of AI-generated AD. By focusing on low-risk factual TV genres, this project seeks to pave the way for scalable and reliable AI-generated AD solutions that clearly address the needs of blind and partially sighted audiences.

Aims and objectives

The overall aim is to test the feasibility of using AI models to automatically generate AD for a selected media/broadcast genre, to evaluate the AI outputs in terms of their accuracy, intelligibility, coherence and robustness (safety), and subsequently to develop initial recommendations for further AI development in this area.

Specific objectives

  1. Identify a broadcast/media genre and a data set for the feasibility study; the genre and data set should meet the criteria for low risk; the likely focus will be a set of TV entertainment shows.
  2. Identify 3-4 recent AI models developed to generate verbal descriptions from images or video scenes; apply them to the selected data to assess their capability in producing AD. 
  3. Evaluate the AI outputs using both automatic evaluation methods and evaluation by human linguistic experts; a set of evaluation criteria will be developed for the human evaluation, informed by AD guidelines and previous research on both human and AI-generated AD.
  4. Consult with professional audio describers to gain insight into their strategies for describing the selected genre, providing an additional benchmark for evaluating the AI descriptions.
  5. Gather feedback from blind and partially sighted users on the AI-generated output to evaluate how well it meets their needs and expectation. 
  6. Triangulate the results, develop conclusions and initial recommendations, with specific reference to the analysed genre and implications for other genres, and with reference to ethical implications of AD automation.

For enquiries or potential collaboration on this topic please contact Prof Sabine Braun, the Principal Investigator of the project.

See other research projects carried out at the Centre for Translation Studies.

Research themes

Find out more about our research at Surrey: