Introducing MilliVid, our new method for long-context video generation! MilliVid creates videos that are consistent over long time spans, without using retrieval heuristics or 3D maps! (1/n)
davidcharatan.com/millivid/#
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
I was a guest on the @MIT_CSAIL podcast to discuss robots, AI, and what the near future might look like. @klgiven did a great job steering the conversation, and I think much of our chat is accessible to non-experts - hope you find it interesting!
bit.ly/4zruZdt
Also on
In January, I started "building something new" with an incredible team. Today I finally get to share some first details about what we've been building. We've called it Walden Robotics (waldenrobotics.com).
I thought long and hard about my own reasons for starting this
We are open-sourcing a fine-tuned video policy + pre-trained IDM! In our paper, we demonstrate that this paradigm has the potential for plug-and-play manipulation across embodiments - very exciting :)
Robot learning is moving beyond policies built for one robot, one scene, one task.
At MIT, we’re exploring a different path: turning video world models into embodiment-agnostic robot policies.
Introducing VERA: a 14B video-to-action system that controls robots across