My research focuses on building self-improving intelligence systems for generalizable multimodal perception and reasoning, within the broader context of representation learning for scalable and compositional understanding.
I'm also interested in unified large-scale models for understanding and
generation, and scalable agentic AI systems.
Sep 03, 2026–Our paper, S3T (Self-Supervised Self-Distillation over Time), the first fully self-contained temporal self-distillation framework for continuous video state tracking, is now on arXiv!
Aug 11, 2026–CVPD (Contrastive Counterfactual Visual Process Distillation) is accepted to BMVC 2026 and our preprint is up on arXiv! See y'all in Lancaster 🇬🇧
Jun 2026–VISE (Visual Invariance Self-Evolution) is accepted to ECCV 2026 and our preprint is up on arXiv! See you in Sweden 🇸🇪
Jun 2026–Our paper, Ask, Solve, Generate, a self-evolving framework for unified image understanding and generation, is now on arXiv!
Dec 12, 2024–Proud to have been selected as a recipient of the Sir C. V. Raman Award by
VIT Chennai for my research!
Jun 23, 2024–
I presented our paper
on attention-fused deep CNNs at ICRAS 2024 in Tokyo, Japan!
Selected Publications
Hover over publications for quick
preview
Temporal Self-Distillation: Learning Visual State Tracking in
Videos Without Supervision
arXiv
The first fully self-contained framework for continuous video state tracking: through temporal self-distillation, a denser view of a clip teaches a sparse-view student with identical weights — no labels, no separate teacher, and no added inference cost.
We introduce S3T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model generates its own target, so training requires no labels, separate teacher, or reward signal, and adds no inference cost. On LLaVA-OneVision-2-8B, S3T improves VSTAT accuracy by +1.74 as a single model, +2.38 with souping, and +2.70 with additional vision-encoder adaptation, while prior self-evolving methods leave state tracking largely unchanged. The capability learned from unlabeled synthetic clips transfers to real videos, improving performance by +7.95 on VSTAT-YouTube state-tracking questions and +4.50 on MVBench Action Count.
BibTeX:
@misc{venkatraman2026temporalselfdistillationlearningvisual,
title={Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision},
author={Shravan Venkatraman and Wenshuai Zhao and Mohammad Hassan Vali and Arno Solin},
year={2026},
eprint={2609.04203},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.04203},
}
Perception Before Supervision: Self-Contained Visual
Distillation from Counterfactual Blind Spots
BMVC 2026
Regions where zooming in sharpens the model's own answer reveal
its visual blind spots, which become dense token-level supervision for self-distillation — no
teacher model, tools, or annotations.
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce CVPD (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs. CVPD identifies visual blind spots where zooming into a region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently utilize under full-image conditioning. We propose a three-gate Counterfactual Criterion that identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation. On Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolving baselines across twelve benchmarks, including methods that rely on external GPT-4o supervision, without a single regression. It achieves gains of +3.60 on OCRBench, +3.38 on MMStar Fine-Grained Perception, and +3.08 on MMStar Logical Reasoning, while maintaining or improving performance on broader multimodal benchmarks.
BibTeX:
@inproceedings{venkatraman2026cvpd,
title = {Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots},
author = {Venkatraman, Shravan and Thawakar, Omkar and Thawkar, Ritesh and
Shaker, Abdelrahman and Anwer, Rao Muhammad},
booktitle = {British Machine Vision Conference (BMVC)},
year = {2026}
}
VISE: Paying More Attention to Visual Tokens in Self-Evolving
Large Multimodal Models
ECCV 2026
Geometric and semantic invariance rewards strengthen visual
conditioning in self-evolving multimodal models, with no labels or external rewards.
Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existing self-evolving LMMs optimize answer agreement without ensuring the decoder attends to visual content, relying instead on statistical language priors to produce self consistent outputs. This leads to a persistent failure mode we term visual under-conditioning, where the decoder relies on language priors rather than the image during generation, manifesting as insufficient attention to visual tokens. As a result, current self-evolving LMMs struggle on vision--language understanding tasks such as image captioning and visual question answering. To address this, we propose VISE (Visual Invariance Self-Evolution), a purely unsupervised self-evolving framework that directly regularizes the model's visual conditioning policy through two complementary invariance-based rewards: a geometric invariance reward that enforces spatial consistency under known transformations, and a semantic invariance reward that penalizes evidence-agnostic generation by requiring the model to recognize the absence of evidence when predicted regions are perturbed. VISE operates within a single model without specialist roles, external reward models, or annotations, and is trained on raw unlabeled images. Experiments on 18 benchmarks demonstrate the efficacy of our approach. Using Qwen3-VL-2B as the base model, VISE achieves gains of +16.85 CIDEr on COCO and +19.66 CIDEr on TextCaps, reduces object hallucination by 5.0 Chair-I points, and generalizes across four model families and scales.
BibTeX:
@inproceedings{venkatraman2026vise,
title = {Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models},
author = {Venkatraman, Shravan and Thawkar, Ritesh and Thawakar, Omkar and
Anwer, Rao Muhammad and Cholakkal, Hisham and Khan, Salman and Khan, Fahad Shahbaz},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}