Aller au contenu

Research Engineer - AI Workload & Systems

    • Markham, Ontario
  • a4tku

Job description

Huawei Canada has an immediate 12-month contract opening for a Researcher.

About the team: 

The Computing Data Application Acceleration Lab aims to create a leading global data analytics platform organized into three specialized teams using innovative programming technologies. This team focuses on full-stack innovations, including software-hardware co-design and optimizing data efficiency at both the storage and runtime layers. This team also develops next-generation GPU architecture for gaming, cloud rendering, VR/AR, and Metaverse applications.

One of the goals of this lab are to enhance algorithm performance and training efficiency across industries, fostering long-term competitiveness.

About the job:

Frontier AI Technology Research

  • Track the evolution of state-of-the-art AI model architectures, including Large Language Models (LLMs), Vision Language Models (VLMs), advanced attention mechanisms, and Mixture-of-Experts (MoE) architectures.

  • Analyze the computational characteristics of emerging model architectures and Agentic AI training and inference workloads.

Hardware-Oriented Workload Analysis

  • Develop a systematic framework to map AI applications and workloads to hardware requirements.

  • Identify key computational patterns, design efficient inference deployment strategies, and build analytical performance models.

  • Evaluate the impact of algorithmic and hardware innovations on system performance, and provide quantitative, explainable architectural recommendations for next-generation AI accelerators.

The total target annual compensation (based on 2,080 hours per year) for this position ranges from $127,000 to $225,000 depending on education, experience, and demonstrated expertise.

Job requirements

About the ideal candidate:

  • Strong understanding of modern AI model architectures and emerging trends, including sparse attention, linear attention, Mixture-of-Experts (MoE), and related techniques.

  • Solid knowledge of AI accelerator architectures (e.g., GPUs, TPUs), with deep understanding of memory hierarchy, interconnect technologies, and hardware performance bottlenecks.

  • Hands-on experience with AI kernel development using technologies such as Triton, TileLang, CUDA, or equivalent. Strong understanding of FlashAttention and other state-of-the-art kernel optimization techniques.

  • Experience with modern LLM inference systems and optimizations, including vLLM, SGLang, or similar serving frameworks. Familiarity with the internals of deep learning frameworks such as PyTorch and JAX.

  • Experience using hardware performance analysis and profiling tools, with the ability to develop Roofline-based performance analysis methodologies and tools.

  • Ph.D. in Artificial Intelligence, Computer Architecture, Computer Systems, or a closely related field is an asset.

  • Demonstrated research contributions through publications or influential open-source projects in AI infrastructure, systems, or computer architecture is an asset.

  • Experience deploying and optimizing large-scale AI training or inference systems in production environments is an asset.

Additional Information:

Huawei Canada is committed to a fair, inclusive, and accessible recruitment process. If you require accommodation during any stage of the hiring process, please let us know and we will work with you to meet your needs.

All applications for this position are reviewed directly by our hiring team, we do not use artificial intelligence tools to screen or select candidates.

or