Andrea Vedaldi, Changan Chen, David Novotny, Roman Shapovalov, Natalia Neverova, Filippos Kokkinos, Ben Graham, Ignacio Rocco, Yanir Kleiman, Meta Austin
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.
We introduce Replay, a collection of multi-view, multimodal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as wearable action cameras, and recorded with a large array of microphones at different positions in the room. Overall, the dataset contains over 4000 minutes of footage and over 7 million timestamped high-resolution frames annotated with camera poses and partially with foreground masks. The Replay dataset has many potential applications, such as novelview synthesis, 3D reconstruction, novel-view acoustic synthesis, human body and face analysis, and training generative models. We provide a benchmark for training and evaluating novel-view synthesis, with two scenarios of different difficulty. Finally, we evaluate several baseline state-of-theart methods on the new benchmark.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2307.12067")
get_code_for_paper("2307.12067")
have("2307.12067")
Connect an agent — have() is free.