X / Twitter
🚨CVPR 2025 (Highlight) Paper Alert 🚨
➡️Paper Title: The Scene Language: Representing Scenes with Programs, Words, and Embeddings
🌟Few pointers from the paper
🎯Authors of this paper introduced “The Scene Language”, a visual scene representation that concisely and precisely describes the structure, semantics, and identity of visual scenes.
🎯It represents a scene with three key components: a program that specifies the hierarchical and relational structure of entities in the scene, words in natural language that summarize the semantic class of each entity, and embeddings that capture the visual identity of each entity.
🎯This representation can be inferred from pre-trained language models via a training-free inference technique, given text or image inputs.
🎯The resulting scene can be rendered into images using traditional, neural, or hybrid graphics renderers.
🎯Together, this forms a robust, automated system for high-quality 3D and 4D scene generation.
🎯Compared with existing representations like scene graphs, their proposed Scene Language generates complex scenes with higher fidelity, while explicitly modeling the scene structures to enable precise control and editing.
🏢Organization: @Stanford , @UCBerkeley
🧙Paper Authors: @zhang_yunzhi , @zizhang_li , Matt Zhou, @elliottszwu , @jiajunwu_cs
📝 Read the Full Paper here: https://t.co/V6Y0HXefSX
🗂️ Project Page: https://t.co/VzD1BJhDYT
🧑💻 Code: https://t.co/Bi69eet90g
🎥 Be sure to watch the attached Demo Video - Sound on 🔊🔊
Find this Valuable 💎 ?
♻️QT and teach your network something new
Follow me 👣, @NaveenManwani17 , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements.
#CVPR2025 #highlight
Introducing SoccerAgent from Hugging Face
A multi-agent system for comprehensive soccer understanding, leveraging a new large-scale knowledge base & benchmark! https://t.co/JMtI8ZjwXO
MLX-Audio v0.2.0 is here!
And it’s our biggest update yet 🚀
What’s new:
Speech-to-Speech 💬
- We now have modular STS pipeline. You can speak with your computer by mixing and matching your favourite STT, LM (VLM soon) and TTS models. Yes, you can interrupt it :)
MLX-Audio Swift 📱💻
- We now support MLX Swift to enable developers can build across Mac, IPad and IPhone.
Speech-to-Text 📘
- We now support STT. Starting with Whisper and Parakeet, with more coming soon.
Text-to-Speech 🗣️
- Added SparkTTS
- Improved Dia KVCache and input management
Thank you very much and shout out to the amazing mlx-audio contributors @lllucas, @BenLumenDigital and @senstella01 ❤️
What an amazing release!
Please leave us a star ⭐️
https://t.co/muDYzy10FA
uniform random noise + band-pass filter = https://t.co/GWcdbG2USw
Paper: https://t.co/ZxCKPEQvEh
Code: https://t.co/2QrM3XjAlx https://t.co/UlXtdJtL9g
🎉 Thrilled to share our CVPR 2025 Award Candidate & Oral paper:
🔹 GlobustVP
Convex Relaxation for Robust Vanishing Point Estimation in Manhattan World
🚀 A globally optimal & outlier-robust method for vanishing point (VP) estimation
🧱 Global optimality
💥 Tolerates up to 70% outliers
⚡ No learning, fast runtime
📄 Paper: https://t.co/egcLbjs6dV
💻 Code: https://t.co/5QeHQrHpT6
1/
Agentic RAG attempts to solve this.
The following visual depicts how it differs from traditional RAG.
The core idea is to introduce agentic behaviors at each stage of RAG. https://t.co/ss3fNINcOw
Game On: join us for a special event on May 6th, 5pm Zürich time
We came to play
https://t.co/dVPsYiwqfW https://t.co/RtJlOIREx1
Recently, LightOn released GTE-ModernColBERT, a multi-vector embedding model for semantic search.
They've now evaluated it on LongEmbed, a benchmark for searching in huge documents, and the performance is state of the art by a mile.
Details 🧵 https://t.co/zaLCqAJqMK
UIGEN-T2 is specifically designed to generate HTML and Tailwind CSS code for web interfaces 👀
https://t.co/3CF2h9JlZM
@bdgrabinski How soon til it’s 140% tariff against the makers of The Apprentice movie from last year?