HaNDF: Object-Conditioned Geometric Neural Hand Distance Fields

NeurIPS 2026

Zuhao Liu, Melih Darcan, Zhengdi Yu, Riza Alp Guler, Tolga Birdal
Imperial College London
Noisy hand poses descend the learned distance field toward plausible grasps on a camera and a bottle; side panels show generation, completion and in-the-wild recovery.

Left: HaNDF is a neural distance field fΦ(θ | z) over hand poses θ, conditioned on an object code z. Its zero level-set S(z) is the set of plausible grasps of that object. A noisy pose θ0 is moved onto S(z) by geodesic gradient steps and ends at θT. Right: the same field is used for grasp generation, completion of partially observed hands, and hand recovery from images.

Abstract

We propose, HaNDF, object-conditional neural distance fields as data-driven priors for modeling the plausible hand–object configurations, in the product manifold of hand articulations. While recent unconditional generative priors showed success in modeling human pose, hands are versatile articulated bodies often found in interaction with other objects. This makes it crucial to develop a generative prior which can model the plausible hand poses conditioned on the object under interaction. To this end, our HaNDF proposes a geometry-aware, object-conditioned neural distance field, representing the plausible articulations in the zero-level set of a conditional field. We leverage tools from Riemannian geometry to (i) represent plausible hands in the zero level-set of an object-induced neural field, and (ii) project a given pose onto the this field. Using HaNDF, we can also optimize for a plausible condition, determining the object under interaction. Our extensive evaluations demonstrate that our HaNDF provides a flexible and powerful prior for hand–object interaction, suitable for applications such as pose refinement, reconstruction under occlusion, and physically plausible manipulation synthesis.

Method

Pose representation

Each hand pose is a point on a product manifold of unit quaternions, MH = ℍ1K, with one quaternion per joint. The distance field is learned with respect to the geodesic distance on this manifold.

theta is the set of K unit quaternions q_i, an element of the hand manifold M_H, the K-fold product of unit quaternions
Unit quaternion sphere with its tangent plane, a geodesic, and the exponential and logarithm maps A hand grasping a cylinder with per-joint rotation distributions drawn around each joint

Left: the unit-quaternion sphere ℍ1, the tangent space at q, the geodesic from q to p, and the maps Expq and Logq. Right: per-joint rotation distributions of one grasp.

Object-conditioned distance field

fΦ(θ | z) is a network that maps a pose θ and an object latent z to the geodesic distance from θ to the nearest plausible pose for that object. The plausible poses form the zero level-set S(z), which changes with z.

S(z) is the set of poses theta in M_H with f_Phi(theta given z) equal to zero; f_Phi maps M_H times Z to the non-negative reals

Projection onto the zero-level set uses Riemannian gradient descent on the field. Each step follows the Riemannian gradient, with a step length proportional to the predicted distance. The exponential map returns the updated pose to the manifold, so every iterate remains a valid hand pose.

theta at step t plus 1 equals the exponential map at theta t of minus alpha times f_Phi times the Riemannian gradient of f_Phi divided by its norm plus epsilon
A hand pose is projected step by step along the distance field toward a grasp on a pink camera
Projection of a noisy pose θ0 onto S(z1) for a camera; θT is the pose after T steps.

Optimization with the prior

Denoising, completion, generation and image-based recovery share the objective below: a task-specific data term plus the squared field value as a prior, minimized over the hand pose.

minimize over theta in M_H and z in Z the data loss plus lambda_prior times f_Phi squared plus lambda_z times a regularizer on z

Results

Pose denoising

The slider shows the pose at t = 0, 100, 300, and 1000 projection steps, with the ground-truth pose on the right.

Noisy hand pose next to a bowl.

Noisy pose

Projection step of a hand pose grasping a bowl
t = 0
Ground-truth grasp of the bowl.

Ground truth

Noisy hand pose next to a mug.

Noisy pose

Projection step of a hand pose grasping a mug
t = 0
Ground-truth grasp of the mug.

Ground truth

Denoising with different priors. Columns: ground truth, HMP, H-NRDF (NRDF trained on hand poses), HaNDF.

Qualitative denoising comparison on a knife and a toy car: ground truth, HMP, H-NRDF and HaNDF

In-the-wild hand recovery

HaNDF is added as the hand prior to EasyHOI and AlignSDF; evaluation on HOI4D, OakInk and DexYCB. For hand-focused comparison, the object ground truth is provided during inference. MPVPE and MPJPE: mean per-vertex and per-joint position error of the hand. I.V.: hand–object intersection volume. Lower is better.

HOI4D OakInk DexYCB
Method MPVPE ↓MPJPE ↓I.V. ↓ MPVPE ↓MPJPE ↓I.V. ↓ MPVPE ↓MPJPE ↓I.V. ↓
iHOI19.4220.1241.4414.7515.171.6629.8530.731.26
EasyHOI16.6417.0613.2813.1713.480.9329.6430.082.61
EasyHOI + HaNDF13.9814.3311.4012.8513.120.8928.3128.782.43
AlignSDF10.4010.605.3910.3410.500.8216.8217.190.99
AlignSDF + HaNDF10.2210.434.399.9810.170.8716.5916.970.98

Left: baseline reconstruction. Right: reconstruction with HaNDF. Circles mark missing hand–object contact, object penetration, and implausible hand rotation.

Input photo: a hand holding a knife over a cutting board.

Input

EasyHOI reconstruction of the knife grasp EasyHOI plus HaNDF reconstruction of the knife grasp
EasyHOI + HaNDF
AlignSDF reconstruction of the knife grasp AlignSDF plus HaNDF reconstruction of the knife grasp
AlignSDF + HaNDF
Ground-truth hand and knife.

Ground truth

Input photo: a hand holding a bottle on a table.

Input

EasyHOI reconstruction of the bottle grasp EasyHOI plus HaNDF reconstruction of the bottle grasp
EasyHOI + HaNDF
AlignSDF reconstruction of the bottle grasp AlignSDF plus HaNDF reconstruction of the bottle grasp
AlignSDF + HaNDF
Ground-truth hand and bottle.

Ground truth

Acknowledgements

We thank the authors of the following works for providing their valuable code and data:

NRDF: Neural Riemannian Distance Fields for Learning Articulated Pose Priors
PoseNDF: Modeling Human Pose Manifolds with Neural Distance Fields
EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild
AlignSDF: Pose-Aligned Signed Distance Fields for Hand-Object Reconstruction
HMP: Hand Motion Priors for Pose and Shape Estimation from Video
iHOI: What's in your hands? 3D Reconstruction of Generic Objects in Hands
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction
DexYCB: A Benchmark for Capturing Hand Grasping of Objects
OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object Interaction

T. Birdal was supported by a UKRI Future Leaders Fellowship [grant number MR/Y018818/1]. The authors acknowledge support from the UK AI Research Resource (AIRR Isambard AI) through grant 0251-4584-0945-1 - TopoFound.

BibTeX

@inproceedings{liu2026handf,
  title     = {{HaNDF}: Object-Conditioned Geometric Neural Hand Distance Fields},
  author    = {Liu, Zuhao and Darcan, Melih and Yu, Zhengdi and Guler, Riza Alp and Birdal, Tolga},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}