Senior AI Engineer
Unknown · Remote (Remote) · posted Sep 15, 2026
Open to candidates worldwide
More remote jobs open to candidates in India, Vietnam, Indonesia, Pakistan, Nigeria, Bangladesh, Malaysia, Philippines, Brazil and Colombia
What this role actually asks for
Extracted by RemoteHuntMust have
- •6+ years in software engineering (ML/AI)
- •Deep Python proficiency
- •Orchestrating coding agents
- •Model Context Protocol (MCP) mechanics
- •LLM evaluation pipelines
- •Open-weight model families
Nice to have
- •Mixture of Experts (MoE)
- •low latency
- •Spanish
- •Portuguese
- •Japanese
Tools and technologies
Languages required
The full posting
About the Role
Design and orchestrate multi-agent workflows, managing planning, implementation, and independent validation agents. Architect MCP integrations and determine optimal modularization strategies for async, heavy-lifting workloads. Establish and maintain spec-driven (SDD) and eval-driven (EDD) development workflows, including written constraint files to prevent agent drift. Build robust LLM evaluation pipelines using LLM-as-judge setups, confusion matrices, and recall/precision tracking. Evaluate, fine-tune, and deploy open-weight models across various hosting and inference providers. Contribute to large-scale document intelligence pipelines, focusing on OCR, classification, and named entity recognition (NER). Collaborate with technical leadership to translate complex regulatory and business logic into reliable agent workflows.
Required Skills & Experience
- 6+ years of experience in software engineering, with a strong focus on shipping production-grade machine learning and AI systems.
- C1 English level or higher.
- Deep proficiency in Python, including advanced data structures and algorithms.
- Hands-on experience orchestrating coding agents (such as Claude Code or Codex) within real engineering workflows.
- Strong working knowledge of Model Context Protocol (MCP) mechanics, tool calls, and agent-to-agent (A2A) protocols.
- Experience building and maintaining LLM evaluation pipelines, including golden datasets, LLM-as-judge frameworks, and recall/precision tracking.
- Familiarity with open-weight model families (such as Qwen, Llama, or Mistral) and Mixture-of-Experts (MoE) architectures.
- Experience with major model-hosting and inference platforms, including AWS Bedrock, Vertex AI, Azure OpenAI, Cerebras, or Base10.
- Hands-on experience with Mixture of Experts, low latency.
Skills
- Python
- LLM
- Model Context Protocol
- Claude Code
- Codex
- Machine Learning
- AWS Bedrock
- Vertex AI
- Azure OpenAI
- Qwen
- Llama
- Mistral
- Mixture of Experts
- Claude
- AWS
- Azure OpenAI
- low latency
- OCR
- OpenCV
- AWS Textract
- Amazon SageMaker
- Spanish
- Portuguese
- Japanese
Is this one actually worth your time?
RemoteHunt scores every remote job 0–100 against your own resume, so you apply to the handful that fit instead of the hundred that don't. Free plan, no card required.