XRelevance 90
Introducing GEN-1.5, a one-shot learner.
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the...
Open sourceHF Daily PapersRelevance 90
RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal ...
Open sourceHF Daily PapersRelevance 90
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. ...
Open sourceXRelevance 90
"Variable Form Personal Robot"
"Variable Form Personal Robot" On stairs or uneven terrain, it takes the form of a dog type (four legs), and on flat floor surfaces, it autonomously changes into a two-legged humanoid type with wheels, adapting to the surrounding environment. https://youtu.be/...
Open sourceXRelevance 90
『Open-Source Physical AI Quadrupedal Robot』
『Open-Source Physical AI Quadrupedal Robot』 With a single desktop robot, you can learn robotics, coding, AI, and 3D printing. https://youtu.be/0Z7ZijSYeV8 #PhysicalAI #quadrupedal #OpenSource #RobotKit #programmable #STEM #Quaddle #PetoiCamp
Open sourceXRelevance 90
This is one of the more unusual flying robots I have seen.
This is one of the more unusual flying robots I have seen. Researchers at the University of Tokyo built DRAGON, a transformable aerial robot that can change its shape while flying. Instead of a rigid frame, it is made of four linked segments. Each segment has ...
Open sourceHF Daily PapersRelevance 90
In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing
Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This creates growing demands for effective testing to ensure system functionality and safety. However, ADS testing remains complex and lacks well-estab...
Open sourceHF Daily PapersRelevance 90
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or ...
Open sourceHF Daily PapersRelevance 90
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both aligned with action generation and sufficiently geometry-aware to capture where and...
Open sourceNVIDIARelevance 90
Beyond VLAs: How World Action Models Reshape Robot Manipulation
A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene...
Open sourceNVIDIARelevance 90
Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 Super
Autonomous vehicle (AV) development often relies on separate models for trajectory generation, high-level intent prediction, scene understanding, and data...
Open sourceGitHub GrowthRelevance 90
Robbyant/lingbot-world-v2: +18 GitHub stars
Infinite Worlds with Versatile Interactions
Open sourceHF Daily PapersRelevance 90
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over...
Open sourceHF Daily PapersRelevance 90
Masked Visual Actions for Unified World Modeling
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visua...
Open sourceHF Daily PapersRelevance 90
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynami...
Open sourceHF Daily PapersRelevance 90
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized,...
Open sourceHF Daily PapersRelevance 90
Generative World Renderer at the Speed of Play
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/control-hints prompts, AlayaRenderer preserves scene structure without altering the ...
Open sourcehnRelevance 90
Xiaomi-Robotics-1
Breaking the data barrier. Scaling robot policy models with embodiment-free pre-training. Foundation models in language and vision keep moving the frontier by riding empirical scaling laws: capability tracks data, parameters, and compute. Robotics has missed o...
Open sourceRedditRelevance 90
MiniCPM-Robot model series - MiniCPM-RobotManip & MiniCPM-RobotTrack
🚀 MiniCPM enters the physical world — enabling robots to understand, remember, and act. We open-source MiniCPM-Robot, our first embodied AI model series, including: 🤖 MiniCPM-RobotManip — a 1.5B general-purpose Vision-Language-Action (VLA) model for robotic ma...
Open sourceGitHub GrowthRelevance 90
Robbyant/lingbot-world-v2: +18 GitHub stars
Infinite Worlds with Versatile Interactions
Open sourceHF Daily PapersRelevance 90
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to n...
Open source36KrRelevance 90
李飞飞第一篇「触觉」论文:机器人会盲摸麻将了
嚯!这阵容,谁组的局? 李飞飞 ,ImageNet奠基人,美国三院院士; Trevor Darrell ,伯克利BAIR联合创始人,他的Caffe深度学习框架曾是计算机视觉领域最广泛使用的平台之一; Jitendra Malik ,R-CNN的核心贡献者之一,深刻影响了计算机视觉的发展轨迹,同样美国三院院士; Pieter Abbeel ,深度强化学习先驱,伯克利机器人学习实验室主任、BAIR实验室联合主任。ACM计算奖得主; Ken Goldberg ,机器人学泰斗,30年前就深耕机器人抓取和自动化,IEEE会士...
Open sourceHF Daily PapersRelevance 90
UniVR: Thinking in Visual Space for Unified Visual Reasoning
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pu...
Open sourceHF Daily PapersRelevance 90
BadWAM: When World-Action Models Dream Right but Act Wrong
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of ro...
Open sourceHF Daily PapersRelevance 90
SPEAR: A Simulator for Photorealistic Embodied AI Research
Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by in...
Open sourceHF Daily PapersRelevance 90
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of t...
Open sourceHF Daily PapersRelevance 90
RoboTTT: Context Scaling for Robot Policies
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude be...
Open source36KrRelevance 90
给人形机器人当老师,撑起一个百亿市场
在机器人真正走进工厂和家庭之前,先要有人教会它们怎么“像人一样干活”。 你可能在社交平台刷到过这样一段视频,印度流水线上的工人们像平常一样在分拣、装配,或者缝纫、裁剪,但他们的头顶与手腕处的摄像头会记录下每一次动作细节。这其实就是在为训练人形机器人做数据采集。 在国内,类似的工作也开始下沉到兼职市场。“具身智能数据采集员”的招聘帖密集出现,“日结薪资、居家可做、无学历要求”吸引了大量求职者。有人反复抓取水杯、整理衣物、搬动物品, 成为机器人的“AI教练”。 这背后,是具身智能行业正在遭遇的数据饥渴。人形机器人要从演...
Open sourceHF Daily PapersRelevance 90
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints...
Open sourcebluesky globalRelevance 90
Boston Dynamics tries using ‘robot dogs’ for deliveries
Boston Dynamics tries using ‘robot dogs’ for deliveries See Spot deliver.
Open source36KrRelevance 90
别人拼命造人形机器人,这家公司先给机器狗装了只手
国内机器人公司的脑洞,还是太大了。 现在的具身智能公司,要么拼命把机器人造得更像人,两条腿、两只手、十根手指一样不少;要么沿着四足路线,把机器狗做得越来越像一条聪明、听话的真狗。 维他动力却整出了一个“四不像”: 一只机器狗,背上长出了一只手。 2026年7月,维他动力发布“大头EDU版”。这是一款面向开发者、实验室和企业研发的四足机器人平台,背部预留了机械臂等扩展接口。 只要在拓展接口上装上机械臂,一只后背长手的机器狗就诞生了。 除了造“四不像”之外,维他动力身上还有另一个有意思的反差。 就在两个月前,维他动力刚...
Open sourceNVIDIARelevance 90
Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning...
Open source