AWS Trainium / NKI Kernel Expert
Anyone Ai · Argentina - Fully Remote, Uruguay · posted Sep 15, 2026
Open to candidates in Argentina and Uruguay
What this role actually asks for
Extracted by RemoteHuntMust have
- •2+ years NKI kernel development/optimization
- •Experience with AWS Trainium/Inferentia2
- •Strong understanding of tile-based computation
- •Familiarity with memory hierarchy and DMA
- •Ability to evaluate CUDA -> NKI migrations
- •Experience profiling/optimizing on Trainium
Nice to have
- •Experience with AWS Neuron SDK/Compiler
- •CUDA or Triton kernel development
- •Knowledge of NeuronCore-v2 architecture
- •Experience with FP32, BF16, FP8, INT8
- •Benchmarking on Trn1/Trn2 instances
Tools and technologies
The full posting
Anyone AI is recruiting experienced AWS Trainium / Neuron Kernel Interface (NKI) engineers for a specialized project focused on evaluating and improving kernel development tasks for AI workloads. We’re looking for engineers with hands-on experience building or optimizing NKI kernels on AWS Trainium or Inferentia2 hardware who understand how Trainium’s architecture differs from traditional GPU programming.
What You’ll Work On You’ll review and evaluate technical tasks involving: NKI kernel correctness and Trainium-specific development patterns CUDA → NKI kernel migrations Trainium performance optimization and benchmarking Memory management across SBUF, PSUM, and HBM Tile-based computation and DMA scheduling Cross-platform numerical correctness between CUDA/Triton and NKI Trainium-specific performance bottlenecks and optimization opportunities Technical feedback and quality assessment of kernel implementations The work involves determining whether implementations are not only technically correct, but also idiomatic and optimized for Trainium hardware rather than simply translated from GPU-based approaches.
What We’re Looking For 2+ years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface (NKI) Experience working with AWS Trainium and/or Inferentia2 Strong understanding of: Tile-based computation SBUF / PSUM / HBM memory hierarchy Partition dimension constraints DMA orchestration Trainium-specific optimization techniques Ability to evaluate CUDA → NKI migrations Experience profiling and optimizing workloads on Trainium Understanding of numerical differences across GPU and Trainium backends Strong ability to analyze complex technical implementations and provide clear written feedback Nice to Have Experience with the AWS Neuron SDK or Neuron Compiler CUDA or Triton kernel development experience Knowledge of NeuronCore-v2 architecture Experience with FP32, BF16, FP8, and INT8 workloads Experience benchmarking workloads on Trn1 or Trn2 instances Familiarity with nki.language , @nki.jit , or XLA custom calls Experience with technical evaluation, AI/ML data projects, RLHF, or rubric-based assessment Engagement Work Type: Remote Engagement: Part-time, project-based consulting Focus: AWS Trainium / NKI kernel engineering and technical evaluation This is a strong fit for engineers who have worked deeply with AWS Trainium infrastructure and low-level ML kernel optimization and are interested in applying that expertise to technically challenging AI projects.
Is this one actually worth your time?
RemoteHunt scores every remote job 0–100 against your own resume, so you apply to the handful that fit instead of the hundred that don't. Free plan, no card required.