Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.
We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.
WHAT YOU’LL WORK ON
You’ll work with GPU and accelerator kernel tasks involving:
– Kernel implementation and debugging
– CUDA and Triton optimization
– Translation between kernel frameworks
– Hardware migration
– Operator fusion
– Performance profiling and benchmarking
– Numerical correctness verification
– Compilation and runtime debugging
– Memory hierarchy optimization
– Kernel-level AI workload performance
You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.
WHAT WE’RE LOOKING FOR
– 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
– Strong experience with at least two of the following:
– CUDA
– Triton
– NKI / AWS Neuron
– Pallas / JAX
– Strong understanding of GPU performance optimization
– Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers
– Understanding of:
– Memory bandwidth
– Compute throughput
– GPU occupancy
– Shared memory
– Register pressure
– Memory coalescing
– Bank conflicts
– Strong understanding of floating-point numerical correctness and tolerance thresholds
– Experience debugging kernel compilation and runtime issues
– Ability to distinguish software defects, environment problems, and genuine optimization challenges
RELEVANT EXPERIENCE
Candidates should have experience with several of the following types of work:
– Writing kernels from technical specifications
– Translating kernels between CUDA, Triton, or other frameworks
– Migrating kernels across hardware platforms
– Debugging incorrect kernel implementations
– Optimizing kernel performance
– Fusing multiple operations into optimized kernels
NICE TO HAVE
– Experience across both NVIDIA GPU and custom accelerator ecosystems
– Experience with AWS Trainium, TPU, JAX, or other accelerators
– Compiler engineering experience
– Familiarity with MLIR, XLA, or intermediate representation lowering
– Contributions to GPU or ML kernel libraries
– Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls
– Experience with AI model evaluation, RLHF, or technical benchmark development
WHAT YOU’LL BE RESPONSIBLE FOR
– Reviewing GPU and accelerator kernel implementations for correctness
– Comparing outputs against reference implementations
– Evaluating numerical tolerance thresholds
– Reviewing kernel benchmarks and determining whether comparisons are fair
– Identifying performance bottlenecks and optimization opportunities
– Assessing whether performance targets are realistic given hardware limits
– Reviewing kernel translations and hardware migrations
– Identifying compilation, driver, memory, shape, and runtime issues
– Determining whether technical tasks are genuinely difficult or incorrectly configured
– Providing clear, actionable technical feedback
ENGAGEMENT
Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: GPU kernels, performance engineering, debugging, and technical evaluation
This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.