Scene
Preloaded scene
—
Points
—
High-score
—
Time
Upload your own
Point cloud (.npy, .ply, or .pkl)
Loading…
Drag to rotate · Scroll to zoom
Method comparison
Top-down · same scene · same query · hover cards for explanation
CLIP only
?
CLIP only (baseline)
Lifts 2D patch features into 3D, scores by text similarity. Finds semantically similar-looking regions but ignores surface geometry — a flat table can score as high as a graspable handle.
Geometry only
?
Geometry only
Scores points by local surface shape — planarity, normal consistency, curvature — matched to the query's physical requirements. No visual or language understanding; misses semantic context.
Geometry-aware ours
?
Geometry-aware (ours)
Multiplies CLIP and geometry scores. A point must look right AND be shaped right — suppressing CLIP false positives and producing tighter, physically plausible regions. Best IoU on LASO benchmark.
What to look for: our method produces a tighter yellow cluster on the correct affordance region. CLIP-only spreads scores broadly; geometry-only can miss the right object entirely.