[Paper Notes] EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
EventEgoHands++ reconstructs left and right 3D hand meshes from a head-mounted event camera. Its pipeline first detects each hand as a separate instance, removes most of the background events, and then reconstructs the visible hands with an attention module whose computation changes when two, one, or no hands are detected. The output for each hand contains 20 joints and 778 MANO vertices.
On the synthetic N-HOT3D benchmark, the method lowers MPJPE from 64.83 mm to 43.01 mm relative to the earlier EventEgoHands, a 33.7% reduction. On the new real EEH-R dataset, it reaches 34.18 mm MPJPE and runs at about 40 FPS when both hands are visible. The ablations are unusually clear: instance-level detection supplies most of the local accuracy gain, while cross-hand attention mainly repairs the relative placement of the two hands.
The dataset may be the more durable contribution. EEH-R contains 1,019,716 ground-truth annotations from 85 sequences and eight subjects, covering well-lit and 3.5-lux scenes. It also exposes a hard boundary of synthetic event data: a model trained only on N-HOT3D reaches 137.23 mm MPJPE on EEH-R, versus 34.18 mm when trained on real data. Synthetic pretraining followed by real-data fine-tuning recovers only a small additional gain. My read is therefore that this paper is strongest as a real-data study with a sensible detection-and-reconstruction baseline, rather than evidence that synthetic event simulation has solved data scarcity.
Paper information
“EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset” is by Ryosei Hara, Wataru Ikeda, Masashi Hatano, and Mariko Isogawa from Keio University, with Isogawa also affiliated with JST PRESTO. These notes cover the 19-page arXiv:2609.17189v1, submitted on September 15, 2026, and accepted by IEEE Access.
The project page contains results and videos. The official implementation includes both EventEgoHands++ and the earlier EventEgoHands, along with dataset request links. The extended synthetic dataset is documented in the N-HOT3D repository.
1. Why an egocentric event camera sees too much
An event camera emits asynchronous brightness-change events instead of ordinary intensity frames. Fast finger motion therefore does not create conventional motion blur, and the sensor retains a wide dynamic range in dim scenes. Those properties suit wearable hand tracking. The camera also reacts to head motion, however. In an egocentric view, a small rotation can make edges across the whole kitchen or workspace fire at once, leaving the hand events embedded in a dense moving background.
EventEgoHands++ converts the recent event stream into a two-channel locally normalized event surface (LNES), preserving event polarity and recency in a frame-like tensor $I\in\mathbb{R}^{2\times H\times W}$. This makes standard image backbones available, but the estimation still happens one accumulated event frame at a time.
The previous EventEgoHands used one binary foreground mask. That mask could not identify left versus right, so the downstream network always produced two hands, including frames where only one or neither hand was present. Its fixed cross-attention then exchanged information with a missing-hand feature. The new method treats hand visibility and identity as first-class inputs.
2. Detect, separate, then reconstruct
The first stage is a YOLO26 instance-segmentation model. For every detection it predicts a bounding box, a mask, a left/right label, and a confidence score. The highest-confidence instance on each side is retained, and its mask is multiplied with the event frame to produce separate inputs $I_l$ and $I_r$. A $7\times7$ dilation makes the crop less brittle when segmentation misses a narrow finger boundary.
Each visible hand then passes through a shared ImageNet-pretrained EfficientNetV2-S encoder. Its $1280\times7\times7$ feature map enters Adaptive Attention, followed by a MANO decoder that predicts pose $\theta$, shape $\beta$, translation $t$, and global rotation $R$. The final output is one set of 20 joints and 778 vertices per visible hand.
flowchart TD
A["Raw positive / negative events"] --> B["Two-channel LNES event frame"]
B --> C["YOLO26 hand detector"]
C --> D["Left bbox + mask + confidence"]
C --> E["Right bbox + mask + confidence"]
D --> F["Masked left event frame"]
E --> G["Masked right event frame"]
F --> H["Shared EfficientNetV2-S encoder"]
G --> H
H --> I["Adaptive Attention conditioned on visibility"]
I --> J["MANO decoders"]
J --> K["20 joints + 778 vertices per visible hand"]
This split is a practical answer to background-event overload. The reconstruction backbone does not need to discover the hands while simultaneously estimating their 3D shape. It receives two small, identity-preserving event regions instead.
3. Adaptive Attention changes the computation graph
When both hands are visible, self-attention first refines the spatial structure within each feature map, and bidirectional cross-attention then exchanges information between hands:
\[\bar F_h=\operatorname{SelfAttn}_h(F_h),\quad h\in\{l,r\},\] \[\widehat F_l=\operatorname{CrossAttn}(\bar F_l,\bar F_r), \qquad \widehat F_r=\operatorname{CrossAttn}(\bar F_r,\bar F_l).\]If only one hand is detected, that branch uses self-attention and the cross-attention operation is skipped. If neither hand is detected, the sample is skipped. This is more than an attention mask: all tokens from an absent hand would make a masked softmax ill-defined or inject a meaningless feature, whereas conditional execution avoids creating that representation and saves computation.
The training objective combines joint, inter-hand, vertex, and MANO-parameter losses:
\[\mathcal L_{hand} =2\mathcal L_{joints} +\mathcal L_{interhand} +2\mathcal L_{vertices} +\mathcal L_{MANO}.\]$\mathcal L_{interhand}$ measures the error in left-to-right joint offsets. That term and cross-attention have related jobs: wrist-aligned losses can recover each hand independently, while inter-hand supervision forces the pair to occupy a consistent shared 3D configuration.
4. Two datasets, and a visible simulation gap
The paper expands N-HOT3D and introduces EEH-R.
| Dataset | Source | Subjects | Sequences / duration | Ground truth | Split |
|---|---|---|---|---|---|
| N-HOT3D | HOT3D Aria RGB converted by v2e | 9 | 136 / 4.4 h | 480,120 frames; MANO, masks, boxes | 334,190 train / 83,760 val / 62,170 eval |
| EEH-R | DAVIS346 + MoCap gloves + OptiTrack | 8 | 85 / 2.36 h | 1,019,716 annotations; MANO, partial mask/box labels | 636,433 train / 164,727 val / 218,556 eval |
N-HOT3D projects the original HOT3D MANO meshes into the event-camera image to create masks and boxes. The authors visually inspect the source poses, regenerate failed masks, and use two unseen subjects for evaluation. It is a large synthetic benchmark with strong camera motion.
EEH-R records desk activities in kitchen and workspace scenes. A DAVIS346 supplies events and 30 Hz grayscale reference images; IMU-equipped MoCap gloves and a 16-camera OptiTrack system provide 120 Hz 3D ground truth, including 16 joint positions per hand. Plain fabric gloves cover the sensors. The subject-disjoint evaluation uses two people unseen during training and validation.
Lighting is deliberately split between 457 lux on average and 3.5 lux. For well-lit scenes, SAM3 automatically generates masks on 198,410 grayscale frames. Dark images contain too little contrast for that route, so the authors manually annotate 1,000 event frames. Detector performance in darkness saturates after using roughly 25% of those dark training masks together with the well-lit annotations, suggesting that a small, targeted labeling budget can correct a large illumination shift.
The paper also measures why the two datasets do not match. N-HOT3D has more locomotion and head motion, producing more events per frame. Its simulated events cluster near the beginning of each accumulation window, while real events are distributed much more evenly and include sensor noise. The cross-dataset numbers make that mismatch concrete:
| Training route | Test set | R-AUC ↑ | RR-AUC ↑ | MPJPE ↓ | MPVPE ↓ |
|---|---|---|---|---|---|
| EEH-R | EEH-R | 0.691 | 0.551 | 34.184 mm | 32.163 mm |
| N-HOT3D | EEH-R | 0.079 | 0.038 | 137.229 mm | 126.446 mm |
| N-HOT3D → EEH-R fine-tuning | EEH-R | 0.695 | 0.557 | 34.172 mm | 32.207 mm |
Synthetic pretraining is not useless, but its measured benefit after fine-tuning is tiny. Direct transfer fails by a wide margin.
5. Read the reconstruction metrics together with detection
The paper reports four reconstruction metrics over a 0–100 mm PCK range. R-AUC aligns each hand to its own wrist and measures local hand pose. RR-AUC expresses both hands relative to the right wrist, so it also penalizes an incorrect relationship between them. MPJPE and MPVPE are wrist-aligned mean errors for joints and mesh vertices.
There is an important evaluation condition: all four reconstruction metrics include only hands successfully detected by the Hand Detector. A missed hand does not increase MPJPE or MPVPE; detection is scored separately with mask mAP. The reported reconstruction errors therefore describe pose quality conditional on detection, not end-to-end success across every visible hand.
Main reconstruction results
| Dataset | Method | R-AUC ↑ | RR-AUC ↑ | MPJPE ↓ | MPVPE ↓ |
|---|---|---|---|---|---|
| N-HOT3D | EventHands | 0.253 | 0.190 | 105.62 mm | 97.29 mm |
| N-HOT3D | Ev2Hands | 0.236 | 0.209 | 112.56 mm | 107.00 mm |
| N-HOT3D | EventEgoHands | 0.417 | 0.261 | 64.83 mm | 60.53 mm |
| N-HOT3D | EventEgoHands++ | 0.661 | 0.528 | 43.01 mm | 39.96 mm |
| EEH-R | EventHands | 0.583 | 0.287 | 42.08 mm | 39.29 mm |
| EEH-R | Ev2Hands | 0.502 | 0.250 | 52.02 mm | 48.59 mm |
| EEH-R | EventEgoHands | 0.478 | 0.324 | 56.59 mm | 52.94 mm |
| EEH-R | EventEgoHands++ | 0.691 | 0.551 | 34.18 mm | 32.16 mm |
On N-HOT3D, EventEgoHands++ reduces MPJPE and MPVPE by 33.7% and 34.0% relative to EventEgoHands. On EEH-R, the strongest baseline for wrist-aligned errors is the much simpler EventHands; the new method still cuts MPJPE by 7.90 mm (18.8%) and MPVPE by 7.13 mm (18.1%).
The real-data baseline ordering is informative. EEH-R consists of relatively static desk tasks, the camera-to-hand distance varies little, and hands occupy a large part of the frame. Those conditions resemble a third-person cropped-hand setting, so EventHands can beat the earlier egocentric method on local wrist-aligned errors. EventEgoHands remains stronger on RR-AUC because it models the two-hand relationship. EventEgoHands++ improves both aspects.
Low light affects global hand-to-hand placement more than local articulation. For EventEgoHands++, well-lit versus dark R-AUC is nearly unchanged, 0.693 versus 0.689, while RR-AUC falls from 0.606 to 0.476. MPJPE rises from 32.68 mm to 35.57 mm. The hand shape remains recoverable once detected; deciding where both hands lie relative to one another is less stable in the sparse dark-event regime.
6. The ablations assign distinct jobs to detection and attention
| Hand Detector | Adaptive Attention | R-AUC ↑ | RR-AUC ↑ | MPJPE ↓ | MPVPE ↓ |
|---|---|---|---|---|---|
| — | — | 0.548 | 0.389 | 46.79 mm | 43.79 mm |
| — | ✓ | 0.553 | 0.394 | 47.15 mm | 43.78 mm |
| ✓ | — | 0.652 | 0.469 | 44.01 mm | 41.04 mm |
| ✓ | ✓ | 0.661 | 0.528 | 43.01 mm | 39.96 mm |
Attention alone adds almost nothing when the model still processes the full event frame. The detector is the main source of accuracy: it localizes small hands, preserves side identity, and removes background motion. Once those regions are clean, Adaptive Attention raises RR-AUC from 0.469 to 0.528, a 12.6% relative improvement, with a much smaller change in wrist-aligned error. Self-attention mainly improves each hand; cross-attention mainly improves the relationship between them.
The segmentation study is more nuanced than “YOLO always wins.” On N-HOT3D, instance segmentation raises mAP@50–95 from 0.042 to 0.407 over the old U-Net. On EEH-R, U-Net is slightly better at the loose mAP@50 threshold, 0.928 versus 0.901, while the Hand Detector is slightly better at strict overlap, 0.657 versus 0.655. In dark scenes its mAP@50–95 advantage is much larger, 0.639 versus 0.527. The instance detector earns its place through identity, boxes, visibility, and better strict-overlap behavior, rather than a win on every segmentation number.
Runtime is practical for an online wearable pipeline. Over 100 samples, EventEgoHands++ runs at 39.87 FPS with two hands and 45.42 FPS with one. The single-hand path is faster because cross-attention and the missing branch are skipped. EventHands remains far faster at 415.88 FPS, but with substantially worse reconstruction accuracy.
7. Where the method still fails
The paper shows three recurring failures. An object can occlude a hand strongly enough to cause a missed detection; even after detection, hidden fingertips remain ambiguous because the model represents no object shape or contact. Nearly static hands and heads emit few events, especially in darkness, so the detector can lose the hand. Finally, per-frame predictions jitter over time because the system does not model a sequence.
Those failures point to a slightly awkward fact about the representation. Event cameras provide fine asynchronous timing, yet the current system accumulates events into LNES frames and reconstructs each frame independently. LNES keeps recency inside the window, but no temporal state connects consecutive estimates. A 4D model that tracks hands, objects, and contacts across event time could address sparse intervals and stabilize depth simultaneously.
EEH-R also has a deliberate but narrow capture domain: eight subjects, two laboratory desk scenes, and fabric-covered MoCap gloves. Accurate ground truth comes with a visual appearance shift from bare hands. The dataset does not test outdoor lighting, long-range hands, rapid walking, or unconstrained in-the-wild object use. Its one-million-scale annotation count should therefore be read as dense coverage inside a controlled domain, not broad coverage of wearable hand activity.
8. What I would carry forward
The architectural lesson is simple: in egocentric event vision, separating where and which hand from what 3D pose is a strong baseline. Visibility-conditioned execution is also a good systems choice. It prevents absent-hand features from contaminating a two-hand model and lowers single-hand latency without adding a separate network.
The data lesson is stronger. N-HOT3D is useful for controlled ablations and modest pretraining, yet it does not reproduce the temporal statistics or sensor noise of real events. The 4× jump in MPJPE under direct synthetic-to-real transfer is difficult to explain away as a small calibration issue. Future work should place more weight on event-simulator validation, real-event pretraining, and temporal adaptation than on adding another spatial attention block.
If I were building on this paper, I would use EEH-R and EventEgoHands++ as a reference point for three additions: persistent temporal tracks through sparse-event intervals, explicit object/contact features for occlusion, and an end-to-end metric that counts missed hands together with mesh error. That would move the task from accurate reconstruction on detected crops toward continuous hand understanding in the conditions where an event camera is supposed to matter most.
