arxiv:2601.19898
SaraG
SLMLAH
AI & ML interests
VLM, LLM, LMM
Recent Activity
upvoted a paper about 1 month ago
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding authored a paper 8 months ago
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding updated a dataset 8 months ago
MBZUAI/DuwatBench