Multimodal AI & AI Systems Research Intern
Huawei Switzerland
About this role
About the Team
We are a research team based in Europe, working at the intersection of Generative AI, Video, Efficient AI, and Distributed Systems. We collaborate closely with leading universities and research institutions, as well as engineering teams, to explore and build the next generation of efficient multimodal AI systems. We are looking for highly motivated Master’s and PhD students who are excited about multimodal intelligence and interested in working on challenging research problems with real-world, large-scale AI systems.
What You Will Work OnDepending on your background, research interests, and experience, you may contribute to one or more of the following areas: Multimodal & Generative AIVision-Language Models (VLMs)Video generation and understandingWorld ModelsMultimodal agentsDiffusion and generative modelsEfficient AIModel compression and quantizationSparsity and efficient attention mechanismsInference optimization and accelerationEfficient execution of large multimodal modelsLong-Context & MemoryLong-context modelingKV Cache optimizationMemory systems for large AI modelsEfficient information retrieval and context managementAI Systems & InfrastructureDistributed inference and parallel computingScheduling and resource optimizationGPU/NPU memory optimizationHardware-software co-designScalable AI infrastructureAI for Real-World ApplicationsExplore how large multimodal and generative models can be efficiently deployed in practical, large-scale scenariosDevelop and evaluate solutions that improve model performance, scalability, and efficiencyYou will have the opportunity to identify interesting research problems, prototype new ideas, design and conduct experiments, and evaluate solutions at scale. Depending on the project and research outcomes, there may also be opportunities to publish research papers or contribute to open-source projects. What We Are Looking ForCurrently pursuing a Master’s or PhD degree in Computer Science, Electrical Engineering, Artificial Intelligence, Machine Learning, or a related fieldStrong interest in Multimodal AI, Generative AI, LLMs, Computer Vision, Video, or AI SystemsSolid programming skills in Python and/or C++Familiarity with PyTorch and modern deep learning frameworksStrong analytical, problem-solving, and research skillsAbility and motivation to independently explore and prototype new ideasExperience in one or more of the following areas would be an advantage: Vision-Language Models, Video Generation, or Diffusion ModelsLLM inference and optimizationCUDA, GPU, or NPU programmingDistributed training or inferenceQuantization, sparsity, or efficient attentionLarge-scale AI systemsOpen-source AI projectsAcademic research and publicationsWhat We OfferThe opportunity to work on cutting-edge Multimodal AI, Video Generation, World Models, and AI Systems researchClose collaboration with researchers and engineers from leading universities, research institutions, and industry teamsAccess to large-scale AI models and advanced AI computing platformsOpportunities to publish research and contribute to open-source projectsA highly international research environment in EuropePotential opportunities for continued collaboration, thesis projects, or future positionsWho We Are Looking ForWe are particularly interested in students who are not only excited about making AI models smarter, but who also want to explore: How can we make large multimodal models faster, more scalable, and more efficient?
If you are excited about the future of Multimodal AI, Video Generation, World Models, and large-scale AI Systems, we would love to hear from you.