X / Twitter
DINOv2 meets text at #CVPR 2025! Why choose between high-quality DINO features and CLIP-style vision-language alignment? Pick both with dino.txt 🦖📖
We align frozen DINOv2 features with text captions, obtaining both image-level and patch-level alignment at a minimal cost. [1/N] https://t.co/7BTwLxqXNG
what are your best tips for months 0-3 of a newborn
LayerPeeler: image-to-vectorized layers https://t.co/XL5naEVVSB
After ~6 years of building these types of architectures (starting with BERT, eg see Baleen), I think calling these multi-agent systems is a distraction.
This is just software. Happens to be AI software.
It doesn’t seem so complicated once you internalize it’s just a program. https://t.co/PLit8cBdaq https://t.co/Ocgu9DQmGO
Share some hugs to Alex and try the Pixel3dmm space on @huggingface 🤗
@Gradio link: https://t.co/dRCgtQkI4x https://t.co/i3cNMaXmZ8 https://t.co/WbpFWzG5gv
Excited to share our CVPR 2025 paper on cross-modal space-time correspondence!
We present a method to match pixels across different modalities (RGB-Depth, RGB-Thermal, Photo-Sketch, and cross-style images) — trained entirely using unpaired data and self-supervision.
Our approach learns correspondences through contrastive random walks across visual modalities.
#CVPR2025 (1/6)
I see a lot of people make the same mistakes building agents. So we shared a few of the principles we use
https://t.co/lRNYmeZCcC
CVPR 2025 papers pt. 1 - Gaze-LLE
Gaze-LLE simplifies gaze target estimation by building on top of a frozen DINOv2 visual foundation model; SOTA performance; open source code and model
more papers: https://t.co/1VlLn2BWxl
↓ more https://t.co/DbQTbjEOP5
how to prompt to get gorgeous UI and get 10x more out of Lovable, Bolt, Cursor, v0 (35 min tutorial) https://t.co/46TBqzRWtO