Portrait of Nikita Karaev

Nikita Karaev

Member of Technical Staff · Amazon FAR

I’m building general-purpose robots at Amazon FAR, following the acquisition of Pixelwise AI, where I was a founder. We’re continuing our vision of training robots using internet-scale human data. Previously, I was fortunate to do a PhD at Meta AI and the Visual Geometry Group, University of Oxford, supervised by Christian Rupprecht, Natalia Neverova, and Andrea Vedaldi. Before my PhD, I completed an engineering program at École Polytechnique in beautiful Paris. I also enjoy running and all mountain-related activities 🏔️

News

16 updates
  1. We released Agentic STAR, which reconstructs and tracks arbitrary objects from monocular videos via agentic analysis-by-synthesis.
  2. VGGT-Ω is a Best Paper Finalist at CVPR 2026! Try the HF demo.
  3. CoTracker3 and SpatialTrackerV2 are accepted at ICCV 2025!
  4. 🎉 VGGT won the Best Paper Award at CVPR 2025! 🎉
  5. We released VGGT, a feed-forward neural net that directly predicts all key 3D attributes of a scene. Try our HF demo!
  6. Yuri and I left Meta AI and started Pixelwise AI to unlock training robots on internet-scale data via imitation learning!
  7. We’re releasing CoTracker3, a new point tracking model trained on real data.
  8. CoTracker is accepted at ECCV 2024!
  9. VGGSfM is accepted at CVPR 2024 as a highlight!
  10. CoTracker now supports tracking of 10x more points.
  11. We released CoTracker, a model for tracking any pixel in a video.
  12. My first PhD paper, DynamicStereo, has been accepted at CVPR 2023!
  13. I have started my PhD at Meta AI and Oxford!
  14. We climbed Mount Elbrus, the highest mountain in Europe! 🏔️ 5642m
  15. I have completed the internship at FAIR.
  16. I have started a research internship at Facebook AI Research with Natalia Neverova and Andrea Vedaldi.

Publications

7 papers

VGGT-Ω

CVPR 2026 Best Paper Finalist

Jianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev, Johannes Schönberger, Patrick Labatut, Piotr Bojanowski, David Novotny, Andrea Vedaldi, Christian Rupprecht

Feed-forward 3D reconstruction scales predictably with model and data size. VGGT-Ω substantially improves accuracy, efficiency and capabilities for both static and dynamic scenes, trained at unprecedented scale with a dynamic-scene annotation pipeline and self-supervised learning.