---
title: "AWS Trainium / NKI Kernel Expert"
company: "Anyone AI"
company_url: "https://www.remjobs.works/companies/anyone-ai"
url: "https://www.remjobs.works/job/anyone-ai-aws-trainium-nki-kernel-expert-5f61d726-fdcb-4ffa-9c5f-937ddec043e1"
apply_url: "https://jobs.ashbyhq.com/anyone-ai/55157704-873f-45d7-af63-33bd74489f2a"
workplace: remote
location: "Argentina - Fully Remote"
remote_scope: "Argentina - Fully"
employment_type: contract
seniority: mid
role: software-engineering
region: latin-america
skills: ["aws"]
salary: "$65 per hour"
date_posted: 2026-09-15T14:04:45.547Z
first_seen_by_remjobs: 2026-09-15T19:00:41.359Z
---

# AWS Trainium / NKI Kernel Expert

**Anyone AI** · Argentina - Fully Remote

Salary: $65 per hour

Apply: https://jobs.ashbyhq.com/anyone-ai/55157704-873f-45d7-af63-33bd74489f2a

## About Anyone AI

Impulsa tu carrera en AI. Desarrolla tus habilidades con un entrenamiento intensivo y práctico, dictado por expertos, y accede al mercado de la AI.

## About the role

Anyone AI is recruiting experienced **AWS Trainium / Neuron Kernel Interface (NKI) engineers** for a specialized project focused on evaluating and improving kernel development tasks for AI workloads.

We’re looking for engineers with hands-on experience building or optimizing **NKI kernels on AWS Trainium or Inferentia2 hardware** who understand how Trainium’s architecture differs from traditional GPU programming.

#### What You’ll Work On

You’ll review and evaluate technical tasks involving:

- NKI kernel correctness and Trainium-specific development patterns

- CUDA → NKI kernel migrations

- Trainium performance optimization and benchmarking

- Memory management across SBUF, PSUM, and HBM

- Tile-based computation and DMA scheduling

- Cross-platform numerical correctness between CUDA/Triton and NKI

- Trainium-specific performance bottlenecks and optimization opportunities

- Technical feedback and quality assessment of kernel implementationsThe work involves determining whether implementations are not only technically correct, but also **idiomatic and optimized for Trainium hardware** rather than simply translated from GPU-based approaches.

#### What We’re Looking For

- 2+ years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface (NKI)

- Experience working with AWS Trainium and/or Inferentia2

- Strong understanding of:Tile-based computation

- SBUF / PSUM / HBM memory hierarchy

- Partition dimension constraints

- DMA orchestration

- Trainium-specific optimization techniques

- Ability to evaluate CUDA → NKI migrations

- Experience profiling and optimizing workloads on Trainium

- Understanding of numerical differences across GPU and Trainium backends

- Strong ability to analyze complex technical implementations and provide clear written feedback

#### Nice to Have

- Experience with the AWS Neuron SDK or Neuron Compiler

- CUDA or Triton kernel development experience

- Knowledge of NeuronCore-v2 architecture

- Experience with FP32, BF16, FP8, and INT8 workloads

- Experience benchmarking workloads on Trn1 or Trn2 instances

- Familiarity with `nki.language`, `@nki.jit`, or XLA custom calls

- Experience with technical evaluation, AI/ML data projects, RLHF, or rubric-based assessment

#### Engagement

**Work Type:** Remote
**Engagement:** Part-time, project-based consulting
**Focus:** AWS Trainium / NKI kernel engineering and technical evaluation

This is a strong fit for engineers who have worked deeply with **AWS Trainium infrastructure and low-level ML kernel optimization** and are interested in applying that expertise to technically challenging AI projects.

---

Source: Anyone AI's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/anyone-ai-aws-trainium-nki-kernel-expert-5f61d726-fdcb-4ffa-9c5f-937ddec043e1
