✕
Log inSign up
Vincent Sitzmann
953 posts
@vincesitzmann

Vincent Sitzmann

@vincesitzmann
Building AI that learns by interacting with the world. Associate Professor @ MIT, leading the Scene Representation Group (scenerepresentations.org).
Cambridge, Massachusetts
vincentsitzmann.com
Joined February 2016
329
Following
19.8K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Jun 8
    Introducing MilliVid, our new method for long-context video generation! MilliVid creates videos that are consistent over long time spans, without using retrieval heuristics or 3D maps! (1/n) davidcharatan.com/millivid/#
    Image
    00:00
    11
  • @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Sep 3
    This looks like a really cool product, congrats to the @theworldlabs team!
    @theworldlabs
    World Labs
    @theworldlabs
    Sep 1
    Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
    Image
    00:00
    4
  • @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Aug 25
    I was a guest on the @MIT_CSAIL podcast to discuss robots, AI, and what the near future might look like. @klgiven did a great job steering the conversation, and I think much of our chat is accessible to non-experts - hope you find it interesting! bit.ly/4zruZdt Also on
    Image
    Why AI Still Can't Load the Dishwasher | CSAIL Alliances
    From cap.csail.mit.edu
    1
  • @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Jul 15
    Congrats - this looks super exciting, Russ & team!!
    @RussTedrake
    Russ Tedrake
    @RussTedrake
    Jul 15
    In January, I started "building something new" with an incredible team. Today I finally get to share some first details about what we've been building. We've called it Walden Robotics (waldenrobotics.com). I thought long and hard about my own reasons for starting this
    2
  • @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Jun 23
    We are open-sourcing a fine-tuned video policy + pre-trained IDM! In our paper, we demonstrate that this paradigm has the potential for plug-and-play manipulation across embodiments - very exciting :)
    @sizhe_lester_li
    Lester Li
    @sizhe_lester_li
    Jun 23
    Robot learning is moving beyond policies built for one robot, one scene, one task. At MIT, we’re exploring a different path: turning video world models into embodiment-agnostic robot policies. Introducing VERA: a 14B video-to-action system that controls robots across
    Image
    00:00
    1