x manual globalRelevance 90
Welding, grinding, blasting, painting, assembly und inspection...
Welding, grinding, blasting, painting, assembly und inspection... Bring the robot inside the structure where no normal robot fits. Like what? Ballast tanks. Double hulls. Penstocks. Pressure vessels. Refinery and chemical tanks. Boilers. Offshore structures. N...
Open sourceHF Daily PapersRelevance 90
RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal ...
Open sourceHF Daily PapersRelevance 90
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. ...
Open sourcex manual globalRelevance 90
"Variable Form Personal Robot"
"Variable Form Personal Robot" On stairs or uneven terrain, it takes the form of a dog type (four legs), and on flat floor surfaces, it autonomously changes into a two-legged humanoid type with wheels, adapting to the surrounding environment. https://youtu.be/...
Open sourcex manual globalRelevance 90
『Open-Source Physical AI Quadrupedal Robot』
『Open-Source Physical AI Quadrupedal Robot』 With a single desktop robot, you can learn robotics, coding, AI, and 3D printing. https://youtu.be/0Z7ZijSYeV8 #PhysicalAI #quadrupedal #OpenSource #RobotKit #programmable #STEM #Quaddle #PetoiCamp
Open sourcex manual globalRelevance 90
This is one of the more unusual flying robots I have seen.
This is one of the more unusual flying robots I have seen. Researchers at the University of Tokyo built DRAGON, a transformable aerial robot that can change its shape while flying. Instead of a rigid frame, it is made of four linked segments. Each segment has ...
Open sourceRedditRelevance 90
omlab/VLX-Seek-1.5-10B · Hugging Face
VLX-Seek-1.5-10B VLX-Seek-1.5-10B is the open-source 10B model in the VLX-Seek 1.5 family, designed for fine-grained perception and visual grounding in embodied scenarios. It targets practical settings such as drones, robots, robotic dogs, surveillance cameras...
Open sourceHF Daily PapersRelevance 90
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Mee...
Open sourceHF Daily PapersRelevance 90
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video gener...
Open sourcex manual globalRelevance 90
This Chinese robotics company has been quietly selling humanoid robots to over 60 countries.
This Chinese robotics company has been quietly selling humanoid robots to over 60 countries. Chery's AiMOGA Robotics reported sales of more than 590 humanoid robots in the first half of 2026, with cumulative deliveries reaching 2,000 units across Europe, Middl...
Open sourcex manual globalRelevance 90
a Vietnamese AI engineer just built a real-time system that counts objects moving down a conveyor belt.
a Vietnamese AI engineer just built a real-time system that counts objects moving down a conveyor belt. it runs on ultralytics' objectcounter. train a detector, define the region you care about, and it handles the rest. works for anything on a line, parts, pac...
Open sourcex manual globalRelevance 90
A robot dog you’re actually supposed to take apart.
A robot dog you’re actually supposed to take apart. China’s MirrorMe built Black Panther X as a modular, open quadruped platform rather than another closed robot dog. 200+ components, 12 joints, open interfaces + code, 4 m/s speed, 33 Nm joint torque, and supp...
Open sourceHF Daily PapersRelevance 90
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distr...
Open sourceHugging FaceRelevance 90
nvidia/Alpamayo2-Super
Alpamayo 2 Super is a 34B-parameter foundation model designed to tackle multiple autonomous vehicle (AV) development tasks. It combines a 32B VLM backbone with a 2B diffusion expert.
Open sourceTechCrunchRelevance 90
Moove raises $250M to become the backbone of the robotaxi industry
Moove is scaling up the autonomous vehicle fleet management side of its business and plans to someday own, not just manage, Waymo robotaxis.
Open sourcex manual globalRelevance 90
Xiaomi Robotics-1 a new robot foundation model from @XiaomiTech_
Xiaomi Robotics-1 a new robot foundation model from @XiaomiTech_ - 100K+ hours real world robot data - VLA architecture - Cross-embodiment generalization - Fast adaptation to new tasks
Open sourcex manual globalRelevance 90
Nucleus deploys humanoid robots into a factory in under 90 days
Nucleus emerged from stealth after deploying humanoid robots into a real factory in under 90 days. Its team includes alumni from 1X, Foundation, NEURA, Agile Robots, CERN and ESA.
Open sourcehnRelevance 90
Waymo in Dallas
Starting today, anyone in Dallas can download the Waymo app and hail a fully autonomous ride. From the road — August 4, 2026 Back to blog The Waymo Team August 4, 2026 Share on Twitter Share on LinkedIn Share on Facebook Starting today, anyone in Dallas can do...
Open sourceHF Daily PapersRelevance 90
Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories
Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robotic surface cleaning. Conventional wiping often spreads the stain, while scrubbing provides stronger friction but risks damaging the surface. I...
Open sourcex manual globalRelevance 90
Introducing Alpamayo 2 Super: the frontier open reasoning VLA for autonomous vehicles and robotaxis.
Introducing Alpamayo 2 Super: the frontier open reasoning VLA for autonomous vehicles and robotaxis. 34B parameters. Full-surround awareness. Available for commercial use under under OpenMDW-1.1. Download now on @huggingface → https://nvda.ws/3RPqnfV
Open sourcex manual globalRelevance 90
"Put the watering can in the green bin on the bottom shelf."
"Put the watering can in the green bin on the bottom shelf." Google DeepMind's new Gemini Robotics 2 hears that and a humanoid just does it, walks over, bends, grabs, carries, places. No step-by-step programming. The whole body moving as one motion is the leap...
Open sourceHF Daily PapersRelevance 90
In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing
Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This creates growing demands for effective testing to ensure system functionality and safety. However, ADS testing remains complex and lacks well-estab...
Open sourceHF Daily PapersRelevance 90
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or ...
Open sourceHF Daily PapersRelevance 90
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both aligned with action generation and sufficiently geometry-aware to capture where and...
Open sourceNVIDIARelevance 90
Beyond VLAs: How World Action Models Reshape Robot Manipulation
A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene...
Open sourceNVIDIARelevance 90
Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 Super
Autonomous vehicle (AV) development often relies on separate models for trajectory generation, high-level intent prediction, scene understanding, and data...
Open sourcehnRelevance 90
Gemini Robotics 2 brings whole body intelligence to robots
Gemini Robotics 2 brings whole body intelligence to robots
Open sourceHF Daily PapersRelevance 90
PhiZero: A World Model Built Around Physical Language
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dyna...
Open sourceHF Daily PapersRelevance 90
INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models
Forward latent world models predict how actions change a scene, but recover actions for a desired change only through expensive test-time search. We introduce INTACT (INtent-To-ACTion), an end-to-end JEPA that turns action-labeled, reward-free trajectories int...
Open sourcebluesky globalRelevance 90
Google DeepMind’s new AI model can control a robot’s entire body
Gemini Robotics 2 controls a humanoid robot from ‘feet to fingertips.’ www.theverge.com/tech/973276/... The new model can help robots tie up trash bags and unscrew lightbulbs, too.
Open sourcebluesky globalRelevance 90
For the First Time, Zoox Can Charge People for Rides in Its Steering-Wheel-Free Robotaxis
The temporary exemption from federal safety standards will let Amazon’s Zoox launch a real, paid robotaxi service in Las Vegas. www.wired.com/story/zoox-b... The temporary exemption from federal safety standards will let Amazon’s Zoox launch a real, paid robot...
Open source36KrRelevance 90
CMC资本投资LiberAI,瞄准物理AI和世界模型赛道
CMC资本今日宣布完成对 世界模型与具身智能研发企业LiberAI(北京将闲科技有限公司)新一轮数亿元Pre-A+轮融资的投资 。这是CMC资本在具身智能与世界模型领域的关键布局,也是自设立“CMC AI创意基金”后,在AI前沿技术赛道的又一次落子。 聚焦世界模型,CMC资本关注AI与物理世界交互 随着智能从数字世界向物理世界延伸,世界模型被视为人工智能下一阶段的重要增长极。2026年上半年,具身智能及机器人领域融资事件密集,资金加速向具备明确技术壁垒和数据壁垒的头部项目集中。在技术路线尚未收敛的窗口期,CMC资本...
Open sourcex manual globalRelevance 90
Gemini Robotics ER 2 is our most capable embodied reasoning model designed for physical AI
Gemini Robotics ER 2 is our most capable embodied reasoning model designed for physical AI Built as a high-level brain for robotics, the model connects directly to the Gemini Live API. It processes continuous video streams to track progress, call tools, search...
Open sourcex manual globalRelevance 90
Real-world LiDAR is noisy and surface-dependent. You need your simulated LiDAR to also reflect that.
Real-world LiDAR is noisy and surface-dependent. You need your simulated LiDAR to also reflect that.
Open sourcex manual globalRelevance 90
The human wrist is one of nature's greatest engineering designs.
The human wrist is one of nature's greatest engineering designs. Researchers at IRIM Lab (KOREATECH) developed a compact robotic wrist mechanism that mimics the rolling motion of human joints. The result? Smooth multi-axis movement with fewer mechanical compon...
Open sourcex manual globalRelevance 90
For decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand.
For decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand. Today, we take a major stride toward making that dream a reality: Introducing Gemini Robotics 2 from @GoogleDeepMind, the intelligence layer powering the next generat...
Open sourcetelegramRelevance 90
Italian Startup Builds Robot with Sensor Skin
🤖 Italian Startup Builds Robot with Sensor Skin Italian startup Generative Bionics created the robot Gene.01 covered in sensor skin that tracks touch, temperature, proximity, and force. This lets it detect humans before contact and adjust movements. Gene.01 le...
Open sourcebluesky globalRelevance 90
I Got a Free Meal From a Private Chef—Who Filmed It All to Train Robots
A German startup sent a camera-wearing chef to my apartment. In exchange for a free lunch, I let them record every chop and stir to train future humanoids. www.wired.com/story/i-let-... A German startup sent a camera-wearing chef to my apartment. In exchange f...
Open sourceTechCrunchRelevance 90
US government bans new foreign-made humanoids, robot dogs, and solar inverters, citing risks to national security
The ban largely affects U.S. imports from China, which currently dominates the global market for making humanoid robots and solar inverters.
Open sourceTechCrunchRelevance 90
Zoox clears final federal hurdle to launch paid robotaxi service
Federal safety regulators have given Zoox a temporary exemption that will allow the Amazon-owned autonomous vehicle technology company to charge customers for rides in its custom-built robotaxi.
Open source36KrRelevance 90
时薪1美元的机器人即将到来:四位行业领袖解读未来发展趋势
“每帮他们节省一分钟、一小时,我们就基本赚回了机器人的成本。” “每小时1美元的机器人成本与20–40美元的人力成本之间,正在创造一个巨大的利润空间。” “我十分确信,距离机器人‘硬起飞’已不到十年。机器人建造机器人、数据中心、芯片厂、采矿、精炼,形成自给自足的系统,也可能只需三年。” “Neo正在构建一个类似应用商店的平台,让第三方开发者创建并销售技能。” 近期,知名博主Calacanis在《All-In Podcast》中邀请了ANYbotics、1X、波士顿动力和Agility Robotics的四位领导者展...
Open sourcex manual globalRelevance 90
Robot barbers are rolling out across Chinese cities.
Robot barbers are rolling out across Chinese cities. 3D scan. Fully automated. All for 3 yuan. Was that a precise cut? Writer: Samuel
Open sourcex manual globalRelevance 90
Today, we’re launching Tau’s humanoid cleaning service in San Francisco at $30 per hour.
Today, we’re launching Tau’s humanoid cleaning service in San Francisco at $30 per hour. Access is initially invite-only as we scale operations. If you don’t have an invite yet, join the waitlist at http://tau-robotics.com. All footage is shown at 1× speed. Ea...
Open sourceHF Daily PapersRelevance 90
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting ...
Open sourceTechCrunchRelevance 90
DoorDash is building its own drone delivery business
DoorDash has received FAA approval to operate a commercial drone delivery service in the United States.
Open sourceNVIDIARelevance 90
Developing Healthcare Robotics with GPU-Native Medical Physics Simulation
Unlike autonomous driving or industrial robotics, healthcare robotics can’t rely on internet-scale data collection or unlimited real-world experimentation....
Open sourceHF Daily PapersRelevance 90
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing ...
Open sourceHF Daily PapersRelevance 90
Data Pyramid for Embodied Manipulation
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degre...
Open sourceTechCrunchRelevance 90
Lyft and Baidu enter London’s robotaxi battleground as testing begins
Baidu's Apollo Go autonomous vehicles will be available on Freenow, the mobility network that Lyft acquired in 2025.
Open source36KrRelevance 90
比亚迪人形机器人,8月“上岗”
比亚迪人形机器人项目迎来明确落地节点。 7月28日,比亚迪官方向财联社记者证实,其人形机器人产品将于八月正式亮相。 此前,比亚迪曾在“迪空间|郑州馆”发布一张人形机器人宣传海报,海报称“八月初,有个新朋友,想来认识你”。 受此消息影响,比亚迪A股盘中快速拉升,一度涨超2%;板块方面,人形机器人概念震荡拉升,明新旭腾2连板,上纬新材涨超9%,北自科技、三花智控、拓斯达、力星股份、天奇股份跟涨。 公开信息显示,“迪空间”是比亚迪打造的品牌线下沉浸式体验终端,区别于传统4S店的单一销售功能,该空间主打品牌展示、产品体验、...
Open sourcex manual globalRelevance 90
Introducing Waddle Labs: Claude Code for robots.
Introducing Waddle Labs: Claude Code for robots. Connect our API to your robot and enter a prompt, then our agents write code to achieve the task in 20 minutes. @yiding_song @theWaddleLabs
Open sourcetelegramRelevance 90
Neuralink demos mind-controlled wheelchair
🧠 Neuralink demos mind-controlled wheelchair Neuralink showed participants in clinical trials controlling an electric wheelchair using only their thoughts. An implant reads movement intentions in real time and moves a cursor on a screen to direct the chair. Th...
Open sourceTechCrunchRelevance 90
Enigma raises $70M to make controlling a robot as easy as adjusting the volume
The massive seed round was led by Index Ventures and Ribbit Capital, with participation from Sarah Guo's Conviction Partners.
Open sourceTechCrunchRelevance 90
Are brain waves the next unlock for physical AI?
Forget YouTube videos — frontier physical AI models need multiple camera angles, dense annotation, and, soon, brain wave readings.
Open source36KrRelevance 90
工业和信息化部人形机器人与具身智能标准化技术委员会2026年度全体会议暨“标准周”活动在浙江绍兴召开
工业和信息化部人形机器人与具身智能标准化技术委员会2026年度全体会议暨“标准周”活动在浙江绍兴召开 2026年7月15日—17日,工业和信息化部人形机器人与具身智能标准化技术委员会(以下简称标委会)2026年度全体会议暨“标准周”活动(以下简称活动)在浙江绍兴上虞区举行。工业和信息化部副部长柯吉欣,中国电子学会理事会党委书记、中国人形机器人百人会理事长张峰,标委会主任委员、开放原子开源基金会理事长谢少锋,工业和信息化部科技司副司长甘小斌,中国电子学会副秘书长、标委会副主任委员兼秘书长梁靓,浙江省人民政府副秘书长施...
Open sourceGitHub GrowthRelevance 90
Robbyant/lingbot-world-v2: +18 GitHub stars
Infinite Worlds with Versatile Interactions
Open sourceHF Daily PapersRelevance 90
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over...
Open sourceHF Daily PapersRelevance 90
Masked Visual Actions for Unified World Modeling
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visua...
Open sourceHF Daily PapersRelevance 90
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynami...
Open sourceHF Daily PapersRelevance 90
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized,...
Open sourceHF Daily PapersRelevance 90
Generative World Renderer at the Speed of Play
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/control-hints prompts, AlayaRenderer preserves scene structure without altering the ...
Open sourceHF Daily PapersRelevance 90
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and p...
Open sourceHF Daily PapersRelevance 90
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically al...
Open sourceHF Daily PapersRelevance 90
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open...
Open sourceHugging FaceRelevance 90
nvidia/Cosmos3-Edge
Cosmos 3: Omnimodal World Models for Physical AI Model Collection Code White Paper Website
Open sourceHugging FaceRelevance 90
openbmb/MiniCPM-RobotTrack
A Compact Vision-Language-Action Policy for Embodied target Tracking
Open sourceHugging FaceRelevance 90
openbmb/MiniCPM-RobotManip
MiniCPM-RobotManip is a 1.5B vision-language-action model for embodied manipulation with the following highlights: Generalist Manipulation: A unified 1.5B generalist policy that uses one set of weights across all downstream tasks and outperforms larger models ...
Open sourceTechCrunchRelevance 90
Gritt exits stealth with $34 million for robots to build solar plants — then, everything else
Gritt is coming out of stealth with $34 million and plan to automate the hardest tasks on construction sites.
Open source36KrRelevance 90
WAIC观点:最快两年,迎来机器人“ChatGPT时刻”
想象这样一个场景:你对机器人说“帮我把洗衣机里的衣服拿出来晾晒”,它看了看洗衣机的门,规划手臂的开门轨迹,缓缓打开舱门,发现衣物缠在一起,它停顿思考一下,重新调整策略,将缠绕的衣物一件件抽出,整理好挂在晾衣杆上。 这个场景之所以令人兴奋,不是因为机器人“做到了”,而是因为它“想到了”。它是在理解世界、预测后果、调整行动。 这就是世界模型想要实现的东西:让机器像人一样,在行动之前先在脑海里推演一遍“我这么做,是否符合这个世界的物理规律”。 2026年被视为具身智能部署元年,在2026世界人工智能大会上,具身智能企业不...
Open sourceNVIDIARelevance 90
Integrate NVIDIA Omniverse RTX Sensor Simulation Into Existing Apps
Developers building 3D, design, simulation, robotics, and industrial digital twin applications need ways to bring physical AI capabilities into the tools and...
Open sourcex manual globalRelevance 90
Introducing Cosmos 3 Edge: our open frontier world model built to run on-device.
Introducing Cosmos 3 Edge: our open frontier world model built to run on-device. Cosmos 3 Edge helps robots learn and act, autonomous vehicles understand road scenes and predict intent, and vision AI agents reason across live video for smart infrastructure. Wi...
Open sourcex manual globalRelevance 90
Xiaomi-Robotics-1 just dropped on Hugging Face 🔥
Xiaomi-Robotics-1 just dropped on Hugging Face 🔥 A robot foundation model trained on 100,000 hours of real-world manipulation. They turned it loose in a real apartment: folding laundry, loading the washer, doing the dishes, packing a suitcase. Fully autonomous...
Open sourcehnRelevance 90
Xiaomi-Robotics-1
Breaking the data barrier. Scaling robot policy models with embodiment-free pre-training. Foundation models in language and vision keep moving the frontier by riding empirical scaling laws: capability tracks data, parameters, and compute. Robotics has missed o...
Open sourceRedditRelevance 90
MiniCPM-Robot model series - MiniCPM-RobotManip & MiniCPM-RobotTrack
🚀 MiniCPM enters the physical world — enabling robots to understand, remember, and act. We open-source MiniCPM-Robot, our first embodied AI model series, including: 🤖 MiniCPM-RobotManip — a 1.5B general-purpose Vision-Language-Action (VLA) model for robotic ma...
Open sourceGitHub GrowthRelevance 90
Robbyant/lingbot-world-v2: +18 GitHub stars
Infinite Worlds with Versatile Interactions
Open sourceHF Daily PapersRelevance 90
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to n...
Open source36KrRelevance 90
李飞飞第一篇「触觉」论文:机器人会盲摸麻将了
嚯!这阵容,谁组的局? 李飞飞 ,ImageNet奠基人,美国三院院士; Trevor Darrell ,伯克利BAIR联合创始人,他的Caffe深度学习框架曾是计算机视觉领域最广泛使用的平台之一; Jitendra Malik ,R-CNN的核心贡献者之一,深刻影响了计算机视觉的发展轨迹,同样美国三院院士; Pieter Abbeel ,深度强化学习先驱,伯克利机器人学习实验室主任、BAIR实验室联合主任。ACM计算奖得主; Ken Goldberg ,机器人学泰斗,30年前就深耕机器人抓取和自动化,IEEE会士...
Open sourceHF Daily PapersRelevance 90
UniVR: Thinking in Visual Space for Unified Visual Reasoning
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pu...
Open sourceHF Daily PapersRelevance 90
BadWAM: When World-Action Models Dream Right but Act Wrong
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of ro...
Open sourceHF Daily PapersRelevance 90
SPEAR: A Simulator for Photorealistic Embodied AI Research
Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by in...
Open sourceHF Daily PapersRelevance 90
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of t...
Open sourceHF Daily PapersRelevance 90
RoboTTT: Context Scaling for Robot Policies
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude be...
Open source36KrRelevance 90
给人形机器人当老师,撑起一个百亿市场
在机器人真正走进工厂和家庭之前,先要有人教会它们怎么“像人一样干活”。 你可能在社交平台刷到过这样一段视频,印度流水线上的工人们像平常一样在分拣、装配,或者缝纫、裁剪,但他们的头顶与手腕处的摄像头会记录下每一次动作细节。这其实就是在为训练人形机器人做数据采集。 在国内,类似的工作也开始下沉到兼职市场。“具身智能数据采集员”的招聘帖密集出现,“日结薪资、居家可做、无学历要求”吸引了大量求职者。有人反复抓取水杯、整理衣物、搬动物品, 成为机器人的“AI教练”。 这背后,是具身智能行业正在遭遇的数据饥渴。人形机器人要从演...
Open sourceHF Daily PapersRelevance 90
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints...
Open sourcebluesky globalRelevance 90
Boston Dynamics tries using ‘robot dogs’ for deliveries
Boston Dynamics tries using ‘robot dogs’ for deliveries See Spot deliver.
Open source36KrRelevance 90
别人拼命造人形机器人,这家公司先给机器狗装了只手
国内机器人公司的脑洞,还是太大了。 现在的具身智能公司,要么拼命把机器人造得更像人,两条腿、两只手、十根手指一样不少;要么沿着四足路线,把机器狗做得越来越像一条聪明、听话的真狗。 维他动力却整出了一个“四不像”: 一只机器狗,背上长出了一只手。 2026年7月,维他动力发布“大头EDU版”。这是一款面向开发者、实验室和企业研发的四足机器人平台,背部预留了机械臂等扩展接口。 只要在拓展接口上装上机械臂,一只后背长手的机器狗就诞生了。 除了造“四不像”之外,维他动力身上还有另一个有意思的反差。 就在两个月前,维他动力刚...
Open sourceNVIDIARelevance 90
Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning...
Open sourceHF Daily PapersRelevance 90
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack syst...
Open source36KrRelevance 90
物理AI,让自动驾驶迎来了“第二春”
就在昨天,一则消息在科技圈激起涟漪: 字节跳动正探索进入自动驾驶领域,由Seed旗下的世界模型团队负责,首站瞄准无人物流 。官方回应滴水不漏——“物理AI领域有很多早期研究和探索,但并没有做智能驾驶业务的计划。” 翻译一下:这事在做,先别急着按Waymo的故事来估值。 但真正值得玩味的,不是字节要不要“造车”,而是它 带谁进场 。过去十年,自动驾驶的叙事主角一直是工程师和路测驾驶员——攒最多的车队、跑最多的里程、堆最贵的传感器。字节这次带来的,却是一帮研究世界模型的人。 这两种人看自动驾驶的眼光,完全是两个维度的东...
Open source36KrRelevance 90
融资2亿,机器人开始帮农民种地:省水92% 省地94%
一家用机器人和AI种菜的公司,刚拿到3000万美元(2亿元)C轮融资。 这家公司叫 Hippo Harvest ,成立于2019年,总部位于美国加州。本轮融资由北美大型温室运营商Cox Farms领投,Congruent Ventures、Hawthorne Food Ventures等跟投。 这家公司融资有点小猛。2024年2月,公司曾完成2100万美元B轮融资。仅最近两轮,它就获得了至少5100万美元资金。 Hippo Harvest的核心产品是水肥系统、机器人。官方称,自己在温室里种出的有机生菜和菠菜,已经可...
Open sourceHF Daily PapersRelevance 90
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a ...
Open sourceHF Daily PapersRelevance 90
ABot-N1: Toward a General Visual Language Navigation Foundation Model
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations direc...
Open source