Grasp Anything 2D — Checkpoint 10501 (grasp_contact 专用)

基于 NVIDIA LocateAnything-3B 的语言引导二维抓取模型,multigt + Vision LoRA 阶段产物。

任务

grasp_contact — 输出平行夹爪两个接触点 (x1, y1, x2, y2),四坐标均为 [0, 1000] 离散 token。

评测结果(RealVLG 官方 mini633)

Split n gAcc_corrected_strict mIoU_strict center_err_median (px)
seen 243 66.67% 57.52% 10.33
similar 226 50.00% 49.03% 16.55
novel 164 28.05% 26.86% 56.87
  • 评测协议: evaluate_realvlg_contact.py,fast 模式
  • contact 专用 checkpoint 中性能最强

使用方法

from locate_anything_service.model import LocateAnythingRuntime
from locate_anything_service.config import Settings

settings = Settings(model_id="charlesH777/grasp-anything-10501")
runtime = LocateAnythingRuntime(settings)
result = runtime.predict("image.jpg", "抓取红色杯子", mode="grasp_contact")

训练配置

  • 阶段: multigt + Vision LoRA (MoonViT LoRA rank 8, last 4 layers)
  • LLM LoRA: rank 32
  • 训练量: seen_contact_blocks = 168,016
Downloads last month
10
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support