PRITHIVSAKTHIUR/Molmo2-HF-Demo

A Gradio-based demonstration for the AllenAI Molmo2-8B multimodal model, enabling image QA, multi-image pointing, video QA, and temporal tracking. Users upload images or videos, provide natural language prompts.

/ 100

Experimental

No Package No Dependents

Maintenance 6 / 25

Adoption 3 / 25

Maturity 9 / 25

Community 0 / 25

How are scores calculated?

Stars

Forks

—

Language

Python

License

Apache-2.0

Category

multimodal-visual-grounding

Last pushed

Dec 24, 2025

Commits (30d)

GitHub

Multimodal Visual Grounding · 25 tools

Get this data via API

curl "https://pt-edge.onrender.com/api/v1/quality/nlp/PRITHIVSAKTHIUR/Molmo2-HF-Demo"

Open to everyone — 100 requests/day, no key needed. Get a free key for 1,000/day.

Higher-rated alternatives

TheShadow29/awesome-grounding

awesome grounding: A curated list of research papers in visual grounding

microsoft/XPretrain

Multi-modality pre-training

TheShadow29/zsgnet-pytorch

Official implementation of ICCV19 oral paper Zero-Shot grounding of Objects from Natural...

TheShadow29/VidSitu

[CVPR21] Visual Semantic Role Labeling for Video Understanding (https://arxiv.org/abs/2104.00990)

zeyofu/BLINK_Benchmark

This repo contains evaluation code for the paper "BLINK: Multimodal Large Language Models Can...

Explore NLP Tools

All categories Trending NLP directory Insights