X / Twitter
🚀 Just wrapped up my internship at Google with a paper submission!
LODGE is a new large-scale level-of-detail 3DGS method, enabling efficient rendering even on mobile devices.
Thanks to the amazing team at Google Zurich for making this possible!
📄https://t.co/lsleNSb6VN https://t.co/rn4xJCN0as
🪄Introducing Anymate—a large-scale dataset of 230K 3D assets with rigging and skinning annotations!
With this dataset, we trained an auto-rigging model and benchmarked a variety of architectures!
🔥Turn static assets into animatable ones in seconds: https://t.co/77ApgJtia1 https://t.co/9UsxmgCwht
Today, we are introducing Manus slides!
Manus creates stunning, structured presentations—instantly. With a single prompt, Manus generates entire slide decks tailored to your needs. Whether you're presenting in a boardroom, a classroom, or online, Manus ensures your message lands. Want edits? Just click and adjust. Once you are done, export or share it with colleagues and peers
Just found out through about "seriation" through @gwern - basically a goated way of organizing a list when you have a neural embedding for each item https://t.co/YGfq6AC3OA
ElevenLabs just lost the crown - to open source.
Chatterbox by @resembleai (I remember writing about this team back in 2022 as a pioneer of AI audio) just released as an open-source alternative for audio generation and voice cloning.
- Zero-shot voice cloning from just 5 seconds of audio
- Unique emotion intensity control—from subtle to dramatically expressive
- Real-time voice synthesis faster than real-time inference
- Built-in watermarking for secure, trusted audio
- Consistently preferred over ElevenLabs in blind evaluations
Fully open-source. No hidden restrictions. This is the way.
Check out the demo:
.@erikphoel on his 3yo reading like a 9yo.
Similar to our experience. Though we are about 1 year later than he is. And instead of phonics tutoring we did classical Montessori—materially different, but also very structured.
And similar outcomes. Our 5yo reads like a ~10yo. https://t.co/KlNynVCABs
Ezra Klein: Writers who outsource their research to AI operate on a flawed model of how the mind works. https://t.co/QaQ0CjT8h3
Want to recognize a song from just a few seconds of distorted audio?
Use Constellation Maps.
The math is brilliantly simple.
With just a handful of bytes; discarding 99% of the waveform, you can recognize a unique fingerprint across hundreds of millions of tracks. https://t.co/BDjeEPiT4X
MAC-VO: Metrics-Aware Covariance for Learning-based Stereo Visual Odometry
TL;DR: learning-based stereo; learned metrics-aware matching uncertainty for dual purposes: selecting keypoint and weighing the residual in pose graph optimization. https://t.co/OmugOAYZXH
I've been building voice agents for the last 6mo and I think the chat-supervisor pattern is a game changer.
Stitched model (STT-LLM-TTS) is slow, but realtime audio models aren't (yet) as smart as text.
This has the best of both worlds. Here's how it works: https://t.co/HUfTUqKF72
We're excited to release v1.0.0 of the Google Gen AI TypeScript/JavaScript SDK!
Some highlights:
- Easily build apps powered by Gemini 2.5 models
- Live API support
- MCP support
- TTS models
- Image & video generation models
- Gemini API & Vertex support
Happy building! https://t.co/E274RGtxx2
🎶Introducing Moonbeam🌕: a foundation model for symbolic music, pretrained on 81.6K+ hours of diverse MIDI data across different instruments, genres and formats (e.g., performance and score MIDI) https://t.co/R8DT9SsYBB
重なった部分を「泡」のように変形する表現
#Illustrator https://t.co/UdTGhlBFIY