Lead the architecture and implementation of large-scale, production-grade distributed systems powering AI features.
Own architecture decisions, distributed-systems strategy, service and pipeline reliability, and technical debt across the AI Products estate.
Build interfaces and operational harnesses that safely transition machine learning models from research into production.
Design scalable infrastructure for generative AI features and integrate production-ready Large Language Model technologies into product features.
Operate production services and pipelines, including on-call participation, incident response, and long-term service ownership.
Mentor junior engineers, champion engineering excellence, and collaborate with scientists, engineers, product managers, analysts, and teams across Xero.
Requirements
At least 5 years of experience building and operating production Python or equivalent-language services at scale.
Strong software engineering, system design, and coding proficiency.
Experience owning production services or pipelines, including operational responsibility, on-call work, incident response, and technical debt management.
Deep understanding of distributed processing principles and strong SQL capabilities.
Experience integrating machine learning models or LLM-based features into production systems.
Familiarity with MLFlow, TensorFlow, PyTorch, Airflow, or Prefect is valued.
Experience applying or fine-tuning LLMs in a product context is a preferred qualification.
Benefits
Flexible hybrid work model in Toronto with office days and collaborative boost days.
CAD $185,000–$225,000 annual base salary range, plus eligibility for annual bonus and equity programs for permanent employees.
World-class health, wellness, and retirement programs.
Wellbeing days, generous leave, and dedicated professional development budgets.
Xero
Trusted by 5M around the world on the most loved SMB accounting platform. Xero's Community Guidelines: https://www.xero.com/support/community-guidelines/