【文章标题】:Launch HN: Hebbian Robotics (YC S26) – Build scalable robotics data pipelines 【HN首发】Hebbian Robotics(YC S26)——构建可扩展的机器人数据管道
【文章正文】: Open source SDK for scalable multimodal data pipelines in robotics and physical AI 面向机器人学与实体AI的可扩展多模态数据管道的开源SDK
Hebbian Robotics (YC S26) is building HFlow, an open source SDK for scalable multimodal data pipelines in robotics and physical AI. It makes data tooling and practices typically developed inside large robotics teams accessible to teams of any size. Hebbian Robotics(YC S26)正在开发HFlow——一个面向机器人学与实体AI的可扩展多模态数据管道的开源SDK。它将通常由大型机器人团队开发的数据工具和实践,赋能给任何规模的团队。
We believe processing data is a major bottleneck in robotics. A corpus can combine video, state, actions, timestamps, and metadata from many recording systems. Teams often feel the problem first in quality control: determining whether cameras froze, streams drifted out of sync, required topics disappeared, or duplicate recordings entered the corpus. As the corpus grows, fragmented scripts make it difficult to know what ran, audit the results, or reproduce a dataset. 我们认为数据处理是机器人学的核心瓶颈。一个数据集可能包含来自多个记录系统的视频、状态、动作、时间戳和元数据。团队首先会在质量控制环节遇到问题:需要判断摄像头是否冻结、数据流是否失步、关键主题是否丢失或是否存在重复记录。随着数据集增长,碎片化脚本会使得运行追踪、结果审计和数据复现变得困难。
Teams can start with HFlow’s built-in checks, write new transformations, checks, labels, and enrichments, or connect processing code they already use. HFlow handles the orchestration, storage, versioning, and curation around those steps. 团队可以从HFlow的内置检查项起步,编写新的转换规则、检查项、标签和增强功能,或接入现有处理代码。HFlow负责这些步骤的编排、存储、版本管理和数据治理。
HFlow stamps each processed episode with its provenance, renders the pipeline as a graph, and records metadata and quality evidence in a queryable catalog. You can trace how outputs were produced, monitor every stage, and investigate a corpus without loading the underlying recordings. HFlow会为每个处理后的数据片段标注溯源信息,将管道可视化为图谱,并将元数据和质量证据记录在可查询目录中。您可以追溯输出生成路径、监控每个阶段,且无需加载原始记录即可分析数据集。
MCAP is HFlow’s v1 input and output boundary because it efficiently stores and serves synchronized video, state, action, and other time-series streams. That format requirement does not define where the data comes from: human-worn cameras, teleoperated robots, autonomous policies, and other collection systems can all feed the pipeline once their data is represented as a supported MCAP episode. MCAP格式是HFlow v1的输入输出边界,因其能高效存储并提供同步的视频、状态、动作等时间序列数据流。该格式要求并不限定数据来源:穿戴式摄像头、遥操作机器人、自主策略等采集系统,只要数据能转换为支持的MCAP片段,均可接入管道。
Status: pre-v1, with the core lifecycle working end to end. HFlow is ready to try locally. See what is implemented and open issues for current details and remaining work. 当前状态:v1前版本,核心生命周期已实现端到端贯通。HFlow已支持本地试用。具体实现内容和待办事项请参阅开放议题。
Help grow the open robotics community. Star the repository, share it with your network, or contribute. Our goal is an open source community where anyone can participate in building the future of robotics. No robot hardware is required to contribute. 助力开放机器人社区发展。欢迎为仓库加星、分享传播或参与贡献。我们的目标是打造一个人人皆可参与构建机器人未来的开源社区,贡献者无需拥有实体机器人硬件。
| HFlow’s boundary | |
|---|---|
| 输入 | 直接支持标准MCAP片段;通过hflow import lerobot支持LeRobot Dataset v3仓库 |
| 处理 | 您的Python转换/检查/标签/增强代码 |
| 执行 | 开发时进程内运行;生产环境生成Airflow 3 DAG进行调度 |
| 持久化输出 | 规范化MCAP片段、溯源数据、衍生文件和Parquet目录 |
| 治理 | 生成版本锁定的清单的DuckDB SQL |
Human and robot data move through a four-stage lifecycle: collection —> ingestion ---------------> curation ------> delivery (landing (transform -> QC gate -> (SQL over (curated MCAP + bucket) enrich, as an episode manifest; convert Airflow DAG) catalog) for training) 人类与机器人数据经历四阶段生命周期: 采集 —> 摄入 -----------> 治理 —> 交付 (原始数据桶) (转换->质检门限->增强, (基于片段目录的 (治理后的MCAP+ 以Airflow DAG运行) SQL操作) 清单;转换为训练格式)
-
Your processing code stays yours. Transformations, quality checks, labels, and enrichments are plain Python functions in your own environment. Existing code plugs in through small adapters instead of being rewritten for a proprietary framework.
-
您的处理代码始终属于您。转换、质检、标签和增强都是您自有环境中的普通Python函数。现有代码通过小型适配器接入,无需为专有框架重写。
-
Episodes are MCAP, the container that ROS 2 records natively and Foxglove/Rerun open directly, written with two tunings described in Dyna’s article: in-band H.264 with GOP length matched to how the data is read, and topic-group chunking (camera streams and state streams never share a chunk, so a training sample costs one read per group instead of one per topic).
-
数据片段采用MCAP格式——ROS 2原生记录的容器格式,Foxglove/Rerun可直接打开,并采用Dyna文章描述的两种优化:带内H.264编码(GOP长度匹配数据读取方式)和主题组分块(摄像头流与状态流永不共享数据块,使得训练样本每组只需一次读取而非每主题一次)。
-
Processed episodes carry their provenance. The file itself records the schema, pipeline, and tool versions that produced it, plus its source URI when available. Catalog records connect measurements and outcomes to step versions, making it easier to trace a bad result back to its origin.
-
处理后的片段携带溯源信息。文件自身记录生成它的模式、管道和工具版本,以及可获取时的源URI。目录记录将测量值与结果关联到步骤版本,便于追溯问题根源。
-
The pipeline is visible as a graph. HFlow renders Airflow DAGs so you can see how stages connect and monitor task status, logs, retries, and reruns.
-
管道可视化呈现。HFlow渲染Airflow DAG,您可直观查看阶段连接关系,并监控任务状态、日志、重试和重新运行情况。
-
Quality checks produce reusable evidence. Accessors extract the inputs existing processing code expects (numpy arrays, MP4 paths, JPEG frames), and results land as queryable measurements rather than hardcoded verdicts. Different datasets can apply different thresholds without processing the media again.
-
质量检查生成可复用证据。访问器提取现有处理代码所需的输入(numpy数组、MP4路径、JPEG帧),结果以可查询的测量值形式存储而非硬编码判断。不同数据集可应用不同阈值,无需重新处理媒体文件。
-
Query the corpus without loading the recordings. Metadata, quality measurements, tags, version stamps, and artifact locations live in the Parquet catalog. DuckDB can answer corpus-wide questions and build manifests without opening the underlying MCAP files.
-
无需加载记录即可查询数据集。元数据、质量测量值、标签、版本戳和衍生文件位置均存储于Parquet目录。DuckDB能回答全库级问题并构建清单,而无需打开底层MCAP文件。
Open DuckDB’s browser over the catalog at any time, including before the first run starts: hflow catalog ui 随时通过以下命令打开DuckDB浏览器访问目录(包括首次运行前): hflow catalog ui
The open-source deployment is built to be easy to own: run one single-tenant workspace with the included Docker Compose runtime, or deploy its generated DAG bundle into an Airflow 3 environment you already oper 开源部署设计为易于掌控:使用内置Docker Compose运行时启动单租户工作区,或将生成的DAG包部署到现有Airflow 3环境中继续操作