[EXCLUSIVE FOR VINUNI] NVIDIA | AI Research Intern, TAO Multi-Modal Model Development - 2026

Hanoi, HCMC

Full-time

07/08 — 21/08/2026

Job Description

To apply for this job, you need to complete both steps below:

STEP 1:

Please send your CV to this email to submit your application directly to the company: 

careerservice@vinuni.edu.vn

Your application will only be received by Recruiter if submitted via above email.


STEP 2:

Kindly scroll to the bottom of this page and complete the short VinUni Tracking Form.

Filling out this form alone does not count as applying. Kindly remind this form is not part of the company’s application process. It only helps Careers, Alumni, Industry and Development (CAID) Department discover more opportunities and follow up in case of system issues.

 

Our team is seeking to extend the internship of our current AI Research Intern for the TAO (Train, Adapt, Optimize) Multi-Modal Model Development project, recognizing their exceptional performance and strong alignment with the team’s research goals. Their innovative ideas and technical contributions have significantly enhanced our work. Given the rapidly evolving field of multi-modal AI, encompassing vision-language modeling, universal segmentation, and large-scale model training, extending this internship will provide further growth opportunities for the intern while strengthening our team’s capacity to develop scalable, high-impact AI solutions.

Embark on an exciting journey with NVIDIA, a global leader in AI and accelerated computing. As an AI Research Intern focusing on multi-modal AI and vision-language model development within the TAO framework in Hanoi/HCM City, Vietnam, you will be at the forefront of advancing cutting-edge machine learning research. You’ll collaborate with a talented team of engineers and researchers dedicated to developing state-of-the-art deep learning models for tasks such as image segmentation, cross-modal understanding, and universal representation learning. This internship offers a unique opportunity to contribute to next-generation AI systems with real-world impact across industries—from autonomous vehicles to intelligent content understanding.

 

What you'll be doing:

Develop and fine-tune multi-modal AI models using NVIDIA’s TAO Toolkit and deep learning frameworks.
Contribute to the design and implementation of vision-language models (VLMs) and universal segmentation systems.
Conduct experiments and benchmarking to evaluate model accuracy, robustness, and scalability.
Collaborate with cross-functional teams to integrate your research into production-level pipelines and NVIDIA SDKs.
Participate in research discussions, code reviews, and technical documentation to share insights and improve methodologies.

 

What we need to see:

Currently pursuing a degree in Computer Science, Computer Engineering, or a related field.
Proven experience with machine learning, deep learning, or computer vision model development.
Strong Python programming skills and proficiency with PyTorch or similar frameworks.
Solid understanding of neural network architectures, transformers, and multi-modal learning techniques.
Excellent problem-solving abilities, attention to detail, and a collaborative mindset.
Familiarity with vision-language models, image segmentation, or large-scale pretraining is a strong plus.

Application form

Full Name *
Email Address *
Offices
Hanoi
HCMC
College  *
VinUni Email  *
Your Resume *
To attach your Resume, click here to upload from your Computer.
Security code *

Submit