We propose Dimitra++, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we propose a conditional Motion Diffusion Transformer (cMDT) to model facial motion sequences, employing a 3D representation. The cMDT is conditioned on two inputs: a reference facial image, which determines appearance, as well as an audio sequence, which drives the motion. Quantitative and qualitative experiments, as well as a user study on two widely employed datasets, i.e., VoxCeleb2 and CelebV-HQ, suggest that Dimitra++ is able to outperform existing approaches in generating realistic talking heads imparting lip motion, facial expression, and head pose.
@article{chopin2025dimitra,title={{AI} Killed the Video Star: Audio-Driven Diffusion Model for Expressive Talking Head Generation},author={Chopin, Baptiste and Dhamija, Tashvik and Balaji, Pranav and Wang, Yaohui and Dantcheva, Antitza},journal={International Journal of Computer Vision},year={2025},}
2024
T-BIOM
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
Maheswar Bora, Tashvik Dhamija, Shukesh Reddy, and 4 more authors
IEEE Transactions on Biometrics, Behavior, and Identity Science, 2024
Deepfake generation has witnessed remarkable progress, contributing to highly realistic generated images, videos, and audio. While technically intriguing, such progress has raised serious concerns related to the misuse of manipulated media. To mitigate such misuse, robust and reliable deepfake detection is urgently needed. Towards this, we propose a novel network FauxNet, which is based on pre-trained Visual Speech Recognition (VSR) features. By extracting temporal VSR features from videos, we identify and segregate real videos from manipulated ones. The holy grail in this context has to do with zero-shot detection, i.e., generalizable detection, which we focus on in this work. FauxNet consistently outperforms the state-of-the-art in this setting. In addition, FauxNet is able to attribute – distinguish between generation techniques from which the videos stem. Finally, we propose new datasets, referred to as Authentica-Vox and Authentica-HDTF, comprising about 38,000 real and fake videos in total, the latter created with six recent deepfake generation techniques. We provide extensive analysis and results on the Authentica datasets and FaceForensics++, demonstrating the superiority of FauxNet. The Authentica datasets will be made publicly available.
@article{bora2025deepfake,title={Do You See What {I} Say? {G}eneralizable Deepfake Detection based on Visual Speech Recognition},author={Bora, Maheswar and Dhamija, Tashvik and Reddy, Shukesh and Chopin, Baptiste and Balaji, Pranav and Das, Abhijit and Dantcheva, Antitza},journal={IEEE Transactions on Biometrics, Behavior, and Identity Science},year={2024},}
BIOSIG
Local Distributional Smoothing for Noise-invariant Fingerprint Restoration
Indu Joshi, Tashvik Dhamija, Sumantra Dutta Roy, and 1 more author
In BIOSIG 2024 – International Conference of the Biometrics Special Interest Group, 2024 · Oral Presentation
Existing fingerprint restoration models fail to generalize on severely noisy fingerprint regions. To achieve noise-invariant fingerprint restoration, this paper proposes to regularize the fingerprint restoration model by enforcing local distributional smoothing by generating similar output for clean and perturbed fingerprints. Notably, the perturbations are learnt by virtual adversarial training so as to generate the most difficult noise patterns for the fingerprint restoration model. Improved generalization on noisy fingerprints is obtained by the proposed method on two publicly available databases of noisy fingerprints.
@inproceedings{dhamija2024biosig,title={Local Distributional Smoothing for Noise-invariant Fingerprint Restoration},author={Joshi, Indu and Dhamija, Tashvik and {Dutta Roy}, Sumantra and Dantcheva, Antitza},booktitle={BIOSIG 2024 -- International Conference of the Biometrics Special Interest Group},editor={Boutros, Fadi and Damer, Naser and Fang, Minchao and Gomez-Barrero, Marta and Raja, Kiran and Rathgeb, Christian and Sequeira, Ana and Todisco, Massimiliano},series={Lecture Notes in Informatics (LNI)},publisher={Gesellschaft f{\"u}r Informatik},address={Bonn},year={2024},note={Oral Presentation},}
2023
AppIntel
Semantic Segmentation in Medical Images through Transfused Convolution and Transformer Networks
Tashvik Dhamija, Anunay Gupta, Shreyansh Gupta, and 3 more authors
Recent decades have witnessed rapid development in the field of medical image segmentation. Deep learning-based fully convolution neural networks have played a significant role in the development of automated medical image segmentation models. Though immensely effective, such networks only take into account localized features and are unable to capitalize on the global context of medical image. In this paper, two deep learning based models have been proposed namely USegTransformer-P and USegTransformer-S. The proposed models capitalize upon local features and global features by amalgamating the transformer-based encoders and convolution-based encoders to segment medical images with high precision. Both the proposed models deliver promising results, performing better than the previous state of the art models in various segmentation tasks such as Brain tumor, Lung nodules, Skin lesion and Nuclei segmentation. The authors believe that the ability of USegTransformer-P and USegTransformer-S to perform segmentation with high precision could remarkably benefit medical practitioners and radiologists around the world.
@article{dhamija2023semseg,title={Semantic Segmentation in Medical Images through Transfused Convolution and Transformer Networks},author={Dhamija, Tashvik and Gupta, Anunay and Gupta, Shreyansh and Anjum and Katarya, Rahul and Singh, Ghanshyam},journal={Applied Intelligence},volume={53},pages={1132--1148},year={2023},doi={10.1007/s10489-022-03642-w},}
Hackathon
Amazon ML Challenge 2023 — 2nd Place (Nationwide, ~5000 Teams)
Secured 2nd place out of 5000 teams across India in the Amazon ML Challenge 2023. The task was to predict product length from a dataset of 2.2 million products, each with a title, description, bullet points, and product type ID. Finetuned BERT and RoBERTa models with per-category embedding dictionaries in an end-to-end regression setup; ensembled their predictions with rounding to nearest training-set value, achieving a significant improvement over frozen-embedding baselines.
2022
Sensors
Cross-Domain Consistent Fingerprint Denoising
Indu Joshi, Tashvik Dhamija, Rohit Kumar, and 3 more authors
Performance of state-of-the-art fingerprint denoising model on poor quality fingerprints degrades due to cross-domain shift observed between training and testing domains. To address this limitation, we present a cross-domain consistent fingerprint denoising model, which ensures that the output of two fingerprint images with the same ridge structure, however varying contrast and ridge-valley clarity should be similar. Results indicate that the proposed CDC-GAN outperforms state-of-the-art fingerprint denoising algorithms on challenging publicly available poor quality fingerprint databases.
@article{joshi2022crossdomain,title={Cross-Domain Consistent Fingerprint Denoising},author={Joshi, Indu and Dhamija, Tashvik and Kumar, Rohit and Dantcheva, Antitza and {Dutta Roy}, Sumantra and Kalra, Prem Kumar},journal={IEEE Sensors Letters},volume={6},number={8},year={2022},doi={10.1109/LSENS.2022.3193924},}