Frame and prototype machine learning and agentic solutions for ambiguous trust and safety problems.
Design, build, and productionize end-to-end machine learning pipelines covering feature engineering, training, evaluation, and deployment.
Build and improve abuse behavior detection across multiple defenses.
Design, launch, and iterate on AI agents that automate trust decisions, including orchestration, tool interfaces, and guardrails.
Create benchmarks, evaluation harnesses, and instrumentation to measure model and agent decision quality.
Develop specialized trust and safety models and use LLMs and AI agents to accelerate model development.
Write, review, and ship clean, testable code while improving scalability and reliability.
Work with large-scale structured and unstructured data to improve models.
Partner with frontline defense teams to validate solutions through experiments and holdouts and quantify business and operational impact.
Participate in code reviews, design discussions, and cross-team collaboration.
Requirements
5–10 years of industry experience in applied machine learning and a track record of building and productionizing models at scale.
1–2+ years of hands-on experience with LLMs and generative AI, including agentic frameworks, orchestration, and evaluation.
Strong Python programming skills and familiarity with Scala, Java, or equivalent.
Knowledge of machine learning practices including training/serving skew minimization, A/B testing, feature engineering, model selection, and algorithms such as gradient boosted trees, neural networks, transformers, and deep learning.
Experience with TensorFlow, PyTorch, or equivalent machine learning frameworks and tooling.
Experience building data engineering systems and end-to-end machine learning pipelines for batch and real-time use cases.
Experience designing evaluation methodologies for machine learning or LLM systems, including benchmarks, ground truth, offline and online metrics, and calibration.
Exposure to large-scale software architecture, well-designed APIs, high-volume data pipelines, and efficient algorithms.
Experience with test-driven development, incremental delivery, and deployment practices.
Experience with multimodal models is preferred.
Exposure to trust and risk domains such as fraud detection, anomaly detection, identity, or account integrity is preferred.
Bachelor’s, master’s, or PhD in computer science, machine learning, or a related field.
Benefits
US remote eligible, with occasional office or offsite work as agreed with the manager.
May be eligible for bonus, equity, benefits, and Employee Travel Credits.
Salary: $200k - $235k/yr
Airbnb
Airbnb is an American company operating an online marketplace for short- and long-term homestays and experiences. The company acts as a broker and charges a commission from each booking.