X / Twitter
🆕 A 1 hour Masterclass on @aisdk, now free on YouTube for the first time ever!
Covering:
- Fundamentals of @Vercel's ultra popular AI SDK
- Clone your own Deep Research in 30 mins!
https://t.co/FsIIu1ugNK
Thanks to @nicoalbanese10 for the incredible work in rerecording his workshop given all the updates that have happened since NY!
Microsoft quietly released an MCP server that converts any Office file (Powerpoint, Word, Excel) to markdown: markitdown-mcp https://t.co/Jxth7sJ1qB
New work from the Robotics team at @AIatMeta . Want to be able to tell your robot bring you the keys from the table in the living room? Try out Locate 3D!
interactive demo: https://t.co/aS9WPPmhcF
model & code & dataset: https://t.co/oMWc32VrH9 https://t.co/MvcfDZS3dQ
[1/6] Recent models like DUSt3R generalize well across viewpoints, but performance drops on aerial-ground pairs.
At #CVPR2025, we propose AerialMegaDepth (https://t.co/tDGMVXAFa7), a hybrid dataset combining mesh renderings with real ground images (MegaDepth) to bridge this gap.
Thrilled to share UniRig from Tripo @tripoai @VastAIResearch: an AI model for rapid automatic 3D rigging. Rig diverse models—from humans, animals, to sci-fi creatures—in seconds! UniRig has the potential to streamline animation workflows with unparalleled accuracy & versatility. https://t.co/3p22bZbe1O
Gemini 2.5 Pro and Flash now have the ability to return image segmentation masks on command, as base64 encoded PNGs embedded in JSON strings
I vibe coded this interactive tool for exploring this new capability - it costs a fraction of a cent per image https://t.co/taBN4T7fo0
Regist3R: Incremental Registration with Stereo Foundation Model
Sidun Liu, Wenyu Li, Peng Qiao, Yong Dou
tl;dr: DUSt3R with only one regression head; autoregressive training strategy
https://t.co/rhAP3Cm70y https://t.co/7cwj3t3NWo
St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
@HavenFeng, @junyi42, @QianqianWang5, @yufei_ye, Pengcheng Yu, @Michael_J_Black, @trevordarrell, @akanazawa
DUSt3R-like framework
https://t.co/AJxa37YDqQ https://t.co/7mBC5xzD4U
manually de-stimulating kids content by turning the brightness down / turning it to black and white, turning playback speed to 0.75, and playing it in their second language
stop making "memory for AI" so complicated and drowning it in abstractions
the basic primitives are:
1. extractor: LLM that decides what is important to remember from an interaction
2. writer: LLM that writes data so that it is recallable in the future
that's it - 2 LLM calls
ReZero is out on Hugging Face
Enhancing LLM search ability by trying one-more-time https://t.co/Ua8R4YKdia
@coldhealing I think this is already effectively true in many ways
Talking about media that requires longer attention spans has always been a loose class signifier, and a way to jump SES.
Three creative parenting and screens takes:
1. Huge good small bad https://t.co/SjKWR1US6S