press release
Published: 10 September 2026

Surrey and NVIDIA training fix could let AI-generated scenes respond properly to your controls

AI systems that generate video frame by frame could follow a user’s camera commands far more accurately, thanks to a new training method developed by researchers from the University of Surrey and NVIDIA. 

This improvement matters most where someone steers a generated scene rather than simply watching it. That covers video games built on worlds the AI creates, virtual production sets a director can walk a camera around and simulated environments used to train robots.  

Getting an AI-generated world to turn left when it is told has consistently been difficult for the technology, partly because of a flaw in how such models are typically trained.  Fast video models that produce content one frame at a time (called students) are trained by a second, much slower model (teachers) that marks the work after the fact. This marker is typically what is allowed to look at the whole finished clip at once, so it can judge an early frame using knowledge of frames and camera moves that had not yet happened when the faster model produced it. 

The research team call this the teacher–student context mismatch. In reality, the student model is being graded against a standard it can never meet once deployed. 

The team’s method, Context-Matched Distillation, rebuilds the teacher so it can only look backwards, then grades each generated frame against the actual history the student produced during its own trial run rather than a reconstruction. Because early attempts tend to wander off course, the method also adds a controlled amount of noise to that history, so the teacher is not distracted by rough patches in the student’s early work. 

 Tested against seven existing pipelines on standard video generation benchmarks, the approach produced the highest overall quality scores and a substantial improvement in camera accuracy, specifically recording the lowest camera position errors of any method tested, on both the easy and hard test sets. 

Generating roughly 30 seconds of video, a much harder task because small errors accumulate into visible drift, the method scored highest on overall quality while producing more movement than rival systems, several of which achieved stability largely by generating less motion in the first place. 

In a blind comparison where an AI judge was shown pairs of videos without being told which system made them, the researchers’ models were preferred in between 60 and 88 per cent of matchups against each of six competing systems. 

The method was built on NVIDIA’s Cosmos-Predict2.5-2B video model and works for both single-frame and multi-frame generation. The researchers note it also avoids an expensive preparation stage that competing approaches require, and that extending it to longer videos does not force the teacher to process more footage at once. 

The study has been posted as a preprint. 

 ###

Notes to editors 

  • Hmrishav Bandyopadhyay and Professor Yi-Zhe Song are available for interview; please contact mediarelations@surrey.ac.uk to arrange. 
  • The preprint is available at: https://arxiv.org/abs/2608.13391 
  • Project page with video examples: https://hmrishavbandy.github.io/cmd-site/ 
  • The paper is a collaboration between NVIDIA and SketchX in the Surrey Institute for People-Centred AI at the University of Surrey. Full author list: Hmrishav Bandyopadhyay, Xuanchi Ren, Zijian Huang, Jay Zhangjie Wu, Tianshi Cao, Ruilong Li, Bryan Chu, Sanja Fidler, Yi-Zhe Song, Zian Wang. 

Share what you've read?

Media Contacts


External Communications and PR team
Phone: +44 (0)1483 684380 / 688914 / 684378
Email: mediarelations@surrey.ac.uk
Out of hours: +44 (0)7773 479911