Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Principal Software Developer

Home > Javascript/Typescript programming jobs

Principal Software Developer in Canada new

  • Markham

Responsibilities

  • Own ROCm’s end-to-end validation architecture across functional, workload, performance, stress, stability, scale-out, and system-level test layers.
  • Define release-qualification gates, exit criteria, coverage standards, performance baselines, stability targets, scale targets, and RAS criteria.
  • Architect distributed test runners, CI fleets, hardware-lab orchestration, result data lakes, flaky-test detection, bisection automation, and developer pre-submit pipelines.
  • Establish GitHub-based quality workflows, PR gating policies, required checks, code-coverage standards, bug-bash cadences, and issue-management practices.
  • Lead root-cause analysis of complex multi-component, multi-node, hardware, firmware, and customer escalations.
  • Drive validation of multi-GPU server nodes, PCIe, Infinity Fabric, xGMI, BMC/IPMI, thermal and power behavior, firmware interactions, and Ethernet/InfiniBand/UALink fabrics.
  • Lead AI/ML and HPC workload validation for training, inference, recommender systems, scientific kernels, and benchmark suites.
  • Mentor Senior and Staff validation engineers, SDETs, and SQA leads through technical reviews and written guidance.
  • Influence validation roadmaps for next-generation Instinct GPUs and represent ROCm validation in customer, OEM, and open-source engagements.

Requirements

  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related discipline, or equivalent demonstrated experience.
  • Software engineering experience in validation, SDET, or quality engineering, including leadership of complex systems validation.
  • Expert Python skills for test automation and infrastructure and strong C++ skills for debugging and production-code extensions.
  • Deep expertise in at least two relevant areas, including GPU software stacks, AI/ML frameworks, HPC runtimes and communication libraries, Linux kernel or drivers, accelerator firmware, or distributed systems.
  • Experience validating multi-GPU, multi-node server platforms through stress, soak, fault-injection, and RAS testing.
  • Experience defining release-qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Experience contributing to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience with agentic AI workflows, automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters of 256 or more GPUs, including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, HPC benchmark methodologies, performance validation, profiling tools, hardware-lab automation, and pre-silicon or first-silicon accelerator bring-up.

Benefits

  • Hybrid role located in San Jose, California.
  • AMD benefits are offered; details are provided through AMD’s benefits overview.

AMD

We care deeply about transforming lives with AMD technology to enrich our industry, our communities, and the world. Our mission is to build great products that accelerate next-generation computing experiences – the building blocks for the data center, artificial intelligence, PCs, gaming and embe...

Similar positions

Frontend Software Engineer

  • Sumundi
  • Contract
  • Remote
  • 09/23/2026
  • Ghana

Software Engineer, Full-Stack

  • Ema
  • Full time
  • Canada
  • 09/23/2026
  • Salary: $135k - $300k/yr
  • Vancouver

Software Engineer II

  • UiPath
  • Full time
  • Others
  • 09/23/2026
  • Bengaluru, India

Senior Software Engineer - React Native

  • Kraken
  • Full time
  • Remote
  • 09/23/2026
  • Argentina, Brazil

Webflow Developer

  • Vectra
  • Full time
  • Others
  • 09/23/2026
  • Bengaluru, India