Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Senior AI Engineer

Home > Python programming jobs

Senior AI Engineer in Others

  • Ho Chi Minh City, Vietnam

Responsibilities

  • Design and build local LLM serving environments on GPU hardware and select configurations based on VRAM, memory bandwidth, and workload.
  • Install and maintain NVIDIA drivers, CUDA, cuDNN, and inference engines across single- and multi-GPU deployments using tensor and pipeline parallelism.
  • Optimize LLM inference through quantization, pruning, distillation, sparsity, mixed precision, and calibration techniques.
  • Deploy and tune vLLM, TensorRT-LLM, and TGI for high-throughput inference, including KV-cache optimization, continuous batching, speculative decoding, and optimized kernels.
  • Build benchmarking for time to first token, inter-token latency, throughput, GPU utilization, and cost per million tokens.
  • Deploy quantized models with autoscaling, load balancing, observability, quality regression gates, and A/B testing on real traffic.
  • Research and implement novel inference optimization and model compression techniques.

Requirements

  • Bachelor's degree in Computer Science, Software Engineering, or a related field.
  • 5+ years of software engineering experience focused on ML infrastructure, LLM inference, or model optimization.
  • Hands-on experience deploying and serving LLMs on GPU hardware in production.
  • Strong understanding of quantization and model compression, including FP8, FP4, INT8, INT4, GPTQ, AWQ, SmoothQuant, and QAT.
  • Experience with high-throughput inference engines such as vLLM, TensorRT-LLM, TGI, and llama.cpp.
  • Solid understanding of GPU architecture, CUDA, and the memory-bandwidth-bound nature of LLM inference.
  • Familiarity with KV cache, continuous batching, PagedAttention, and speculative decoding.
  • Proficiency in Python; C++ and CUDA familiarity is a strong plus.
  • Experience writing or tuning custom CUDA or Triton kernels is preferred.
  • Experience with multi-GPU and distributed inference is preferred.
  • AWS or Azure architecture certifications are preferred.
  • Experience with CI/CD and cloud production deployment on Azure, AWS, or GCP, including GPU-backed instances, is preferred.

OPSWAT

OPSWAT protects critical infrastructure. Our goal is to eliminate malware and zero-day attacks. We believe that every file and every device pose a threat. Threats must be addressed at all locations at all times—at entry, at exit, and at rest. Our products focus on threat prevention and process cr...

Similar positions

Graduate Software Engineer

  • Synack
  • Full time
  • UK
  • 10/04/2026
  • Remote

Software Engineer - Cloud Images

  • Canonical
  • Full time
  • Remote
  • 10/04/2026
  • EMEA, Americas

GTM Engineer

  • Higgsfield
  • Full time
  • USA
  • 10/04/2026
  • Salary: $130k - $175k/yr
  • San Francisco, CA

Research Engineer

  • Clera
  • Full time
  • Others
  • 10/04/2026
  • Zürich, Switzerland

Lead Automation Quality Engineer

  • London Stock Exchange Group
  • Full time
  • Others
  • 10/04/2026
  • Salary: Competitive
  • Bengaluru, India