-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathfeed.xml
More file actions
249 lines (249 loc) · 13.5 KB
/
Copy pathfeed.xml
File metadata and controls
249 lines (249 loc) · 13.5 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
<title>Awesome Egocentric Atlas updates</title>
<id>https://chaoyue0307.github.io/awesome-egocentric-atlas/feed.xml</id>
<link href="https://chaoyue0307.github.io/awesome-egocentric-atlas/feed.xml" rel="self"/>
<link href="https://chaoyue0307.github.io/awesome-egocentric-atlas/"/>
<updated>2026-08-09T00:00:00Z</updated>
<subtitle>Recently added egocentric AI resources, newest first.</subtitle>
<entry>
<title>VLAff / EgoAffordance</title>
<id>https://ojh6404.github.io/vlaff/</id>
<link href="https://ojh6404.github.io/vlaff/"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>EgoAffordance contains 204K episodes with 5.6M visual affordances and 11.6M grasp and trajectory affordances; VLAff predicts heatmaps, grasp poses, and executable trajectories</summary>
</entry>
<entry>
<title>SoniSpeech</title>
<id>https://doi.org/10.7298/xjjr-9m85</id>
<link href="https://doi.org/10.7298/xjjr-9m85"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>34 hours and 18,000 voiced and silent utterances with synchronized ultrasound echoes, audio, and frontal video; 5,356 unique words and full phoneme coverage</summary>
</entry>
<entry>
<title>SmartRes</title>
<id>https://arxiv.org/abs/2608.01638</id>
<link href="https://arxiv.org/abs/2608.01638"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>Dynamic pixel-space resolution routing for Ego4D and EgoIntention grounding, reducing visual tokens by up to 67%, retaining 86.4% of full-resolution performance, and reporting up to 1.66x faster inference</summary>
</entry>
<entry>
<title>SiMDex</title>
<id>https://lin-nie.github.io/SiMDex/</id>
<link href="https://lin-nie.github.io/SiMDex/"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>Similarity-based mining selects about 1.49M task-relevant samples from a pool of roughly 32M egocentric human samples and raises reported dexterous-manipulation success from 47.7% to 61.1% over equal-sized random sampling</summary>
</entry>
<entry>
<title>RoboReact</title>
<id>https://arxiv.org/abs/2608.03387</id>
<link href="https://arxiv.org/abs/2608.03387"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>Agentic pipeline that generates egocentric RGB-D manipulation video from a single observation, extracts geometry-preserving keyframes, retargets them to humanoids, and refines execution through closed-loop VLM guidance</summary>
</entry>
<entry>
<title>RekaDaily-10k (raw)</title>
<id>https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw</id>
<link href="https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>Incremental raw release currently containing 7,834 hours, 397,171 videos, 9,836 WebDataset shards, and about 70 TB of unscripted first-person daily-life footage from head-mounted and handheld phones</summary>
</entry>
<entry>
<title>ROCO IROS 2026 UMI Dataset</title>
<id>https://huggingface.co/datasets/rocochallenge2025/roco_iros2026_umi_dataset</id>
<link href="https://huggingface.co/datasets/rocochallenge2025/roco_iros2026_umi_dataset"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>4,928 LeRobot v3 episodes and 7,459,440 frames across five tasks at 30 fps, with stereo first-person video, four hand-camera streams, and 16D bimanual pose and gripper state/action</summary>
</entry>
<entry>
<title>Omega-0</title>
<id>https://arxiv.org/abs/2608.06375</id>
<link href="https://arxiv.org/abs/2608.06375"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>Latent predictive whole-body world-action model supporting ego RGB plus exocentric RGB/depth; reports Omega-HOME with 40+ hours of synchronized household humanoid observations, motion, states, and action latents</summary>
</entry>
<entry>
<title>Nexdata 10,000-Hour Egocentric Video Dataset</title>
<id>https://huggingface.co/datasets/Nexdata-AI/10000-Hour-Egocentric-Video-Dataset</id>
<link href="https://huggingface.co/datasets/Nexdata-AI/10000-Hour-Egocentric-Video-Dataset"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>Commercial card describes 10,000 hours across residential, retail, and office scenes with PICO 4 Ultra 4K stereo video, wrist and ankle IMU, 76-point full-body pose, and dense action annotations</summary>
</entry>
<entry>
<title>LAWM-3D</title>
<id>https://arxiv.org/abs/2608.05706</id>
<link href="https://arxiv.org/abs/2608.05706"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>3D-aware latent-action learning from multiview human video using invariant action tokenization, geometric foundation-model alignment, and RGB-D reconstruction for robot world-model pretraining</summary>
</entry>
<entry>
<title>JoyAI-RA 0.5</title>
<id>https://joyai-ra-05.github.io/</id>
<link href="https://joyai-ra-05.github.io/"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>Vision-Language-World-Action framework aligning action-free human egocentric video, simulation, and robot trajectories through latent and canonical action spaces for generalist manipulation</summary>
</entry>
<entry>
<title>HumynLabs Egocentric Sample Collection</title>
<id>https://huggingface.co/humyn-labs</id>
<link href="https://huggingface.co/humyn-labs"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="collection"/>
<summary>Five CC-BY-4.0 sample packs spanning synchronized head and dual-wrist video with IMU, APAC stereo and monocular labels, residential voice-over captions, and LATAM residential IMU</summary>
</entry>
<entry>
<title>HiResNets</title>
<id>https://arxiv.org/abs/2608.02140</id>
<link href="https://arxiv.org/abs/2608.02140"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>Native Full-HD recognition architecture using foveal residual streams and log-polar read/write operations to preserve fine egocentric objects without quadratic feature-grid growth</summary>
</entry>
<entry>
<title>HOPE</title>
<id>https://subin6.github.io/page-hope</id>
<link href="https://subin6.github.io/page-hope"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>Hand-centric video transformer that predicts temporally evolving per-vertex hand pressure and contact by unifying tactile-glove, planar-pressure, and contact supervision</summary>
</entry>
<entry>
<title>GST-Bench</title>
<id>https://arxiv.org/abs/2608.05747</id>
<link href="https://arxiv.org/abs/2608.05747"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="benchmark"/>
<summary>Human-verified global spatial VQA derived from 6,790 minutes of synthetic video, mapping egocentric streams to unseen viewpoints and top-down scenes; the best reported zero-shot VLM scores 42.68 versus 79.08 for humans</summary>
</entry>
<entry>
<title>EventKitchen</title>
<id>https://arxiv.org/abs/2608.04865</id>
<link href="https://arxiv.org/abs/2608.04865"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>5.5 hours from 10 participants in 13 kitchens with egocentric stereo events, synchronized RGB, depth, and IMU, plus 10,762 action segments and 13,482 object boxes</summary>
</entry>
<entry>
<title>EgoSuite-Open10K</title>
<id>https://huggingface.co/datasets/xuejf/EgoSuite-Open10K</id>
<link href="https://huggingface.co/datasets/xuejf/EgoSuite-Open10K"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>Announcement card describes 10,000 selected egocentric and wrist-view clips across six SKUs and seven scene groups, with hand pose, body pose, and semantic annotations</summary>
</entry>
<entry>
<title>EgoAfford</title>
<id>https://egoafford.github.io/</id>
<link href="https://egoafford.github.io/"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="benchmark"/>
<summary>About 15.5K human-verified images from 2,000 generated multi-step scenes plus 102 real egocentric images across 26 tasks, with plans and role-specific affordance masks</summary>
</entry>
<entry>
<title>Ego500</title>
<id>https://huggingface.co/datasets/humanarchive/ego500</id>
<link href="https://huggingface.co/datasets/humanarchive/ego500"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>Card reports 500 hours of first-person work video, more than 310,000 structured action annotations, 390 task-environment combinations, and 60+ environments</summary>
</entry>
<entry>
<title>Ego2Robot</title>
<id>https://www-ye.github.io/ego2robot_blog/</id>
<link href="https://www-ye.github.io/ego2robot_blog/"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>Pipeline converting curated and in-the-wild egocentric human manipulation video into 18,561 hours of robot-format data across 15 robot morphologies through retargeting, visual synthesis, and quality curation</summary>
</entry>
<entry>
<title>DreamTraj / MOVE</title>
<id>https://whathappen0.github.io/DreamTraj/</id>
<link href="https://whathappen0.github.io/DreamTraj/"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="model"/>
<summary>MOVE provides 5,038 object-centric egocentric 6-DoF trajectories with fine-grained language; DreamTraj decodes trajectories from unrendered video-diffusion latents and reports 4.6x faster inference than generate-then-extract pipelines</summary>
</entry>
<entry>
<title>Assistant Placement Aria</title>
<id>https://arxiv.org/abs/2608.00652</id>
<link href="https://arxiv.org/abs/2608.00652"/>
<updated>2026-08-01T00:00:00Z</updated>
<category term="benchmark"/>
<summary>Synthetic and real indoor scenes with 2D images, 3D point clouds, text descriptions, and annotations for panel placement, sitting suggestion, and TV placement</summary>
</entry>
<entry>
<title>yyyyywv/egocentric</title>
<id>https://huggingface.co/datasets/yyyyywv/egocentric</id>
<link href="https://huggingface.co/datasets/yyyyywv/egocentric"/>
<updated>2026-07-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>About 32.8 GB and 1,019 robot episodes totaling 994,459 frames across seven LeRobot task groups: 800 gripper episodes plus plate, shoe, tea, wash, gift-in-hand, and flower-picking sets, with eye or head and wrist cameras, state, and action streams at 20-25 fps</summary>
</entry>
<entry>
<title>Xiaomi-Robotics-U0</title>
<id>https://arxiv.org/abs/2607.11643</id>
<link href="https://arxiv.org/abs/2607.11643"/>
<updated>2026-07-01T00:00:00Z</updated>
<category term="model"/>
<summary>38B unified embodied-synthesis model spanning multi-view scene generation, embodiment transfer, editing, and video generation; reports improving pi0.5 out-of-distribution manipulation success from 36.9% to 63.2%</summary>
</entry>
<entry>
<title>Worldscape-MoE</title>
<id>https://worldscape-moe.com/</id>
<link href="https://worldscape-moe.com/"/>
<updated>2026-07-01T00:00:00Z</updated>
<category term="model"/>
<summary>Mixture-of-experts diffusion world model unifying camera trajectories, robot actions, and hand-joint action maps across locomotion, manipulation, and egocentric hand-control experiments</summary>
</entry>
<entry>
<title>World Action Models to Embodied Brains Roadmap</title>
<id>https://arxiv.org/abs/2607.11689</id>
<link href="https://arxiv.org/abs/2607.11689"/>
<updated>2026-07-01T00:00:00Z</updated>
<category term="survey"/>
<summary>Roadmap organizing WAM gaps across model roles and representations, objectives and standardization, and system composition, then proposing embodied brains, physical harnesses, shared contracts, and closed-loop post-training</summary>
</entry>
<entry>
<title>Whareformer</title>
<id>https://jacobchalk.github.io/Whareformer/</id>
<link href="https://jacobchalk.github.io/Whareformer/"/>
<updated>2026-07-01T00:00:00Z</updated>
<category term="model"/>
<summary>First learning-based solution for online Out of Sight, Not out of Mind tracking, trained on 56 videos and evaluated on 260 long videos from EPIC-KITCHENS-100, IT3DEgo, and HD-EPIC with updatable appearance and 3D-location memory</summary>
</entry>
<entry>
<title>Wearable Gait MoCap with Shank-Mounted Ego Cameras</title>
<id>https://doi.org/10.1038/s41597-026-07657-7</id>
<link href="https://doi.org/10.1038/s41597-026-07657-7"/>
<updated>2026-07-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>Wearable motion-capture dataset for gait analysis using IMUs and shank-mounted egocentric cameras</summary>
</entry>
<entry>
<title>WANDA / Worlds in One Demo</title>
<id>https://wanda.lecar-lab.org/</id>
<link href="https://wanda.lecar-lab.org/"/>
<updated>2026-07-01T00:00:00Z</updated>
<category term="dataset"/>
<summary>Synthetic-data engine expanding one real demonstration per task into 16,922 LeRobot episodes and about 26.8M timesteps (roughly 248 hours at 30 fps) across five long-horizon mobile-manipulation tasks and diverse generated worlds</summary>
</entry>
<entry>
<title>WAM-TTT</title>
<id>https://arxiv.org/abs/2607.06988</id>
<link href="https://arxiv.org/abs/2607.06988"/>
<updated>2026-07-01T00:00:00Z</updated>
<category term="model"/>
<summary>Test-time training framework that steers frozen world-action models from unlabeled raw human videos by adapting a lightweight memory through self-supervised video prediction and paired human-robot meta-training</summary>
</entry>
</feed>