Method Detail: GeoTextCA

Back to Leaderboard
Benchmark: VLMOD
Short name: GeoTextCA
Long name: Geometry-Enhanced Text Cross-Attention Network for Multi-Object Monocular 3D Visual Grounding
Description: We propose a cross-modal matching network for multi-object 3D visual grounding from monocular RGB images. Given an image, three textual descriptions and a set of detected 3D objects, the method predicts for each object which descriptions it matches (a multi-label binary classification over the three descriptions). The method consists of three components: (1) Text encoding. The three textual descriptions are encoded by a pre-trained language model (MiniLM) into semantic sentence embeddings via mean pooling, which are then projected into a common embedding space. (2) Geometry encoding with image-level 3D features. Each object's 16-field structured features (2D bounding box, 3D size, 3D location, orientation, truncation, occlusion, type and color) are normalized and encoded by a multi-layer perceptron with learnable type/color embeddings. In parallel, a convolutional backbone (ResNet-18) extracts image-level 3D features by applying RoIAlign to each object's 2D region, which are fused with the structured geometry features to enhance the object representation. (3) Cross-attention fusion. Object tokens serve as queries and text tokens serve as keys/values in a cross-attention module, followed by a matching head that produces multi-label (N x 3) binary logits indicating which of the three descriptions each object matches. The model is trained with a binary cross-entropy loss over the multi-label matching matrix.
Reference: N/A
Last submitted: September 17, 2026
Published: September 15, 2026 at 04:45:53
Submissions: 12
Project page / code: N/A
Open source: Yes

Benchmark performance

Submission Date F1 (↑) Precision (↑) Recall (↑) TP (↑) FP (↓) FN (↓)
2026-09-17 04:16 70.6202 60.4554 84.8940 2602 1702 463
2026-09-17 03:54 69.9284 59.6908 84.4046 2587 1747 478
2026-09-17 01:18 70.6202 60.4554 84.8940 2602 1702 463
2026-09-15 15:05 66.9725 54.0092 88.1240 2701 2300 364
2026-09-15 12:42 65.2348 54.0523 82.2512 2521 2143 544
2026-09-15 12:38 65.3448 53.7458 83.3279 2554 2198 511
2026-09-15 12:29 64.4258 49.6815 91.6150 2808 2844 257
2026-09-15 12:24 66.9725 54.0092 88.1240 2701 2300 364
2026-09-15 08:37 66.9725 54.0092 88.1240 2701 2300 364
2026-09-15 08:33 63.4916 54.7974 75.4649 2313 1908 752
2026-09-15 08:20 64.9216 52.2249 85.7749 2629 2405 436
2026-09-15 06:02 66.9725 54.0092 88.1240 2701 2300 364