Evaluation of AI-generated audio description for factual TV/media genres
Start date
January 2025End date
OngoingAbout the project
This project aims to evaluate emerging AI solutions, including AI models and AI-based pipelines, to assess their effectiveness in generating simple audio descriptions for a selected factual TV genre. The project will explore the feasibility of leveraging generative AI to make TV content more accessible for blind and partially sighted audiences.
The demand for accessible TV content is growing, driven by increased awareness of diversity among stakeholders and expanding accessibility legislation. However, current accessibility-related practices such as audio description (AD) are costly and time-intensive, as they rely heavily on specialised human expertise. This dependency has led to significant gaps in AD provision, particularly in live TV broadcasts. At present, only around 10% of TV content is audio described, compared to over 90% of TV content that includes subtitles, while both accessibility practices could benefit diverse audiences.
The current gaps, coupled with legislative requirements, have sparked strong interest in AI-driven AD solutions within the media and broadcast sectors. Rapid advances in generative AI hold considerable potential for automating the generation of AD, but current AI technology falls short in key areas. AI-generated AD often lacks the nuance and sophistication of human descriptions, particularly in terms of accuracy and specificity of character, object and action description, in selecting relevant information, and in terms of creating coherent, meaningful narratives across video scenes. While computer vision communities have intensified research into the automatic generation of video descriptions (‘video captioning’), little attention has been paid to the specifics of AD, such as the need to integrate AD with existing film dialogue or narration, and to maintain cohesion between AD fragments and with the overall context of a programme.
Further challenges associated with AI-generated AD include potential bias in training data, which may impact accuracy and user experience, and a lack of transparency in description generation (the black box problem in generative AI), which make it difficult to predict how AI descriptions are generated, potentially resulting in critical errors.
While this means that AI-generated AD is currently unlikely to be suitable for use without human supervision and review, feasibility testing using different AI models and pipelines is crucial at this stage. It will help identify key challenges (e.g. the areas of critical error) and opportunities for safe and systematic development of AI-generated AD. By focusing on low-risk factual TV genres, this project seeks to pave the way for scalable and reliable AI-generated AD solutions that clearly address the needs of blind and partially sighted audiences.
Aims and objectives
The overall aim is to test the feasibility of using AI models to automatically generate AD for a selected media/broadcast genre, to evaluate the AI outputs in terms of their accuracy, intelligibility, coherence and robustness (safety), and subsequently to develop initial recommendations for further AI development in this area.
Specific objectives
- Identify a broadcast/media genre and a data set for the feasibility study; the genre and data set should meet the criteria for low risk; the likely focus will be a set of TV entertainment shows.
- Identify 3-4 recent AI models developed to generate verbal descriptions from images or video scenes; apply them to the selected data to assess their capability in producing AD.
- Evaluate the AI outputs using both automatic evaluation methods and evaluation by human linguistic experts; a set of evaluation criteria will be developed for the human evaluation, informed by AD guidelines and previous research on both human and AI-generated AD.
- Consult with professional audio describers to gain insight into their strategies for describing the selected genre, providing an additional benchmark for evaluating the AI descriptions.
- Gather feedback from blind and partially sighted users on the AI-generated output to evaluate how well it meets their needs and expectation.
- Triangulate the results, develop conclusions and initial recommendations, with specific reference to the analysed genre and implications for other genres, and with reference to ethical implications of AD automation.
People
Professor Sabine Braun
Professor of Translation Studies; Director, Centre for Translation Studies; Co-Director, Surrey Institute for People-Centred AI
Biography
I am a Professor of Translation Studies, Director of the Centre for Translation Studies, and a Co-Director of the Surrey Institute for People-Centred Artificial Intelligence at the University of Surrey in the UK. From 2017 to 2021 I also served as Associate Dean for Research and Innovation in the Faculty of Arts and Social Sciences at the University of Surrey.
My research explores the integration and interaction of human and machine in translation and interpreting, for example to improve access to critical information, media content and vital public services such as healthcare and justice for linguistic-minority populations and other groups/people in need of communication support. My overarching interest lies in the notions of fairness, trust, transparency, and quality in relation to technology use in these contexts.
For over 10 years, I have led a programme of research that has involved cross-disciplinary collaboration with academic and non-academic partners to improve access to justice for linguistically diverse populations. Under this programme, I have investigated the use of video links in legal proceedings involving linguistic-minority participants and interpreters from a variety of theoretical and methodological perspectives. I have led several multi-national research projects in this field (AVIDICUS 1-3, 2008-16) while contributing my expertise in video interpreting to other projects in the justice sector (e.g. QUALITAS, 2012-14, Understanding Justice, 2013-16, VEJ Evaluation, 2018-20). I have advised the European Council Working Party on e-Law (e-Justice) and other justice-sector institutions in the UK and internationally on video interpreting in legal proceedings and have developed guidelines which have been reflected in European Council Recommendation 2015/C 250/01 on ‘Promoting the use of and sharing of best practices on cross-border videoconferencing’.
In other projects I have explored the use of videoconferencing and virtual reality to train users of interpreting services in how to communicate effectively through an interpreter IVY, 2011-3; EVIVA, 2014-15, SHIFT, 2015-18).
A further example of my work on accessibility is my research on audio description (video description) for visually impaired people. In the H2020 project MeMAD (2018-21) I have recently investigated the feasibility of (semi-)automating AD to improve access to media content that is not normally covered by human AD (e.g. social media content).
In 2019, the Research Centre I lead was awarded an ‘Expanding Excellence in England (E3)' grant (2019-24) by Research England to expand our research on human-machine integration in translation and interpreting. As part of this, I am currently leading and involved in a number of pilot studies aimed at better human-machine integration in different modalities of translation and interpreting.
The insights from my research have informed my teaching in interpreting and audiovisual translation on CTS’s MA programmes and the professional training courses that I have delivered (e.g. for the Metropolitan Police Service in London).
From 2018-2021 I was a member of the DIN Working Group on Interpreting Services and Technologies and co-authored the first standard on remote consecutive interpreting worldwide (DIN 8578). I am a member of the BSI Sub-committee Terminology. From 2018-2022, I was the series editor of the IATIS Yearbook (Routledge) and am currently associate series editor for interpreting of Elements in Translation and Interpreting (CUP) and a member of the Advisory Board of Interpreting (Benjamins). I was appointed to the sub-panel for Modern Languages and Linguistics for the Research Excellence Framework REF 2021.
Professor Constantin Orasan
Professor of Language and Translation Technologies
Biography
I am Professor of Language and Translation Technologies at the Centre of Translation Studies, University of Surrey, UK, and a Fellow of the Surrey Institute for People-Centred Artificial Intelligence. Before starting this role, I was Reader (Associate Professor) in Computational Linguistics at the University of Wolverhampton, UK, and the deputy head of the Research Group in Computational Linguistics at the same university. I hold a PhD in computational linguistics and a BSc in computer science.
With over 25 years of experience in the fields of Natural Language Processing, Artificial Intelligence, and Linguistics, I have established myself as a leading researcher in the development of technologies that facilitate access to information. My PhD was in automatic summarisation, and I have led projects on question answering, text simplification, and translation technologies. Notable projects that I have led are EmpASR, an AHRC-funded project focused on training interpreters on how to benefit from the latest developments in artificial intelligence; HarnessingNLP4Court, a UKRI-funded project focused on facilitating access to legal information; the EXPERT project, an Initial Training Network (ITN) funded under the EU’s FP7 to train the next generation of world-class researchers in the field of data-driven translation technology; and the FIRST project, which developed language technologies for making texts more accessible to people with autism.
My current research is interdisciplinary, focusing on the intersection of AI, NLP, and translation studies. In recent years, I have increasingly focused on the practical application of NLP to support translators and interpreters. My recent publications explore reference-less translation evaluation, the processing of multilingual content in low-resource settings, the use of automatic speech recognition to support interpreters, and the use of large language models in text accessibility. My research is well known as a result of over 150 peer-reviewed articles in journals, books, and international conferences.
I am currently leading an EPSRC-funded project focused on making science accessible, and I am Co-Director of the ADA Leverhulme Doctoral Scholarships Network. More information about my work can be found at https://dinel.org.uk/.
Dr Yuan Zou
Lecturer in Translation Studies
Biography
I am a Lecturer in Translation Studies at the University of Surrey's Centre for Translation Studies (CTS). My background spans audiovisual translation (AVT), interpreting, and post-editing. I hold a PhD in AVT from Queen's University Belfast (QUB) and an MTI in Translation and interpreting from Jilin University.
Before joining Surrey, I was teaching Interpreting and Translation at QUB, and I engaged in freelance work as a translator and interpreter. These experiences have been instrumental in shaping my research direction and pedagogical approach.
I am currently focused on the integration of language technologies in the fields of interpreting and audiovisual translation (AVT), with a keen interest in harnessing these advancements to improve digital accessibility. I am actively investigating innovative ways in which technology can be harnessed to support and improve access for individuals with disabilities, ensuring that digital content is more inclusive and accessible to all audiences.
For enquiries or potential collaboration on this topic please contact Prof Sabine Braun, the Principal Investigator of the project.
See other research projects carried out at the Centre for Translation Studies.
Research themes
Find out more about our research at Surrey: