I feel that AI is on the verge of dramatic advances in the general area of video understanding. I really want to develop AI that can interact with humans fluidly via language, understand their actions. We are going to revolutionize the world with the advances that we are going to make. My name is Lorenzo Torresani. I am a full professor in College of Computer Sciences, and I'm also President Aoun Chair. I'm interested in human-centric AI. So, robots and agents that interact with people, communicate with people, and actually support their decision-making. And so I do research in computer vision, specifically in video understanding, which means basically building AI systems that can extract information from video. Most consequential, you know, roles in the real world involve not just analyzing superficially, you know, what is in front of our eyes, but really figuring out why certain actions were taken, right, or performing counterfactual reasoning. So, thinking of what other actions should have been taken if the conditions were different or predicting future outcomes. And so, those are fundamental questions that AI today cannot quite well address. And so my objective is to advance the state of AI so that it can tackle these harder questions and provide beneficial support to humans. SVI-Bench, in the context of video understanding, it goes beyond the problem of determining who is doing what and when in video, and it really tackles the fundamental question of what are the consequences of actions being taken. This is the first benchmark on real-world multi-agent video that tests AI’s capability at four different levels: perception, causal reasoning, counterfactual simulation, and also agency. This is important in the context of AI assistance, for example, where you want your agent to be fully aware of what you are about to do because the assistant must surface relevant, contextual information that is consistent with, not the action you are performing at the present, but what you are going to do next. Team sports actually provide the ideal test bed for all of these questions. They contain lots of interesting human-human interactions. They show coordination and also adversarial behavior. But most importantly, they contain verifiable outcomes. This allows us to easily verify the answers provided by AI, against what truly happened in the right time. [... the right time when to attack...] As you view that video, you fully understand why it succeeded. Jordan took that first aggressive step that caused the defense to overreact, and that opened a pathway that didn't exist before. And that led to the winning jump shot. But if you ask AI today to describe what happened there, they would simply say, “There was a successful jump shot.” They don't understand why the defense collapsed, or they wouldn't be able to tell you what would happen if Jordan decided to drive left instead of taking that aggressive first step to the right. And so, what our research aims to do is to build models that develop these capabilities, understand cause and effect. They can simulate “what if” scenarios so as to support humans in consequential planning. This benchmark has been open source to the entire community. So, we release the raw videos, all of the annotations that we developed, all the training data for pushing the models to actually learn the desired behavior. And it also defines a new protocol for evaluating AI on these challenging tasks. So, a lot of the technology that we developed is meant to be deployed on smart glasses. For example, you know, where you have coordinated efforts among a group of people. For instance, first responders at an emergency scene, imagine them, you know, wearing smart glasses and receiving this assistance automatically by these agents that see the world through the eyes of these first responders and can help them make the right decision, can make them consider plans of coordination that maybe they didn't need vision. So, it is important to develop AI that will actually improve society. For example, a lot of my research is about improving skills – physical skills, like cooking or learning to dance, learning to play a sport. This kind of skill, these days, are difficult to be acquired. Right? So, you either hire a coach, right? But that's expensive. It's only available to a few, or you watch instructional videos, which are very passive, right? Passive way of learning. And so, what I want to do is to build agents that can really interactively help humans to improve their physical skills. And this is a fantastic moment to be an AI scientist, particularly in the field of video understanding. I believe that, you know, we are getting to really challenging problems, but that will have important consequences on society, improve decision-making by humans, and improve, assistance and healthcare, hopefully, the entire world.