---
title: "GPU Kernel Engineer – CUDA, Triton & Accelerator Performance"
company: "Anyone AI"
company_url: "https://www.remjobs.works/companies/anyone-ai"
url: "https://www.remjobs.works/job/anyone-ai-gpu-kernel-engineer-cuda-triton-accelerator-performance-610f982e-8adb-4fe5-9c00-d004510f3f8f"
apply_url: "https://jobs.ashbyhq.com/anyone-ai/c94d1d2a-796e-428f-a8dc-ca41cc144a5b"
workplace: remote
location: "Argentina - Fully Remote"
remote_scope: "Argentina - Fully"
employment_type: contract
seniority: mid
role: software-engineering
region: latin-america
skills: ["aws"]
salary: "$65 per hour"
date_posted: 2026-09-15T14:14:39.369Z
first_seen_by_remjobs: 2026-09-15T19:00:41.359Z
---

# GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

**Anyone AI** · Argentina - Fully Remote

Salary: $65 per hour

Apply: https://jobs.ashbyhq.com/anyone-ai/c94d1d2a-796e-428f-a8dc-ca41cc144a5b

## About Anyone AI

Impulsa tu carrera en AI. Desarrolla tus habilidades con un entrenamiento intensivo y práctico, dictado por expertos, y accede al mercado de la AI.

## About the role

Anyone AI is recruiting experienced **GPU Kernel Engineers** for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as **CUDA, Triton, NKI, or Pallas**, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

#### What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

- Kernel implementation and debugging

- CUDA and Triton optimization

- Translation between kernel frameworks

- Hardware migration

- Operator fusion

- Performance profiling and benchmarking

- Numerical correctness verification

- Compilation and runtime debugging

- Memory hierarchy optimization

- Kernel-level AI workload performanceYou’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

#### What We’re Looking For

- 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels

- Strong experience with at least two of the following:CUDA

- Triton

- NKI / AWS Neuron

- Pallas / JAX

- Strong understanding of GPU performance optimization

- Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers

- Understanding of:Memory bandwidth

- Compute throughput

- GPU occupancy

- Shared memory

- Register pressure

- Memory coalescing

- Bank conflicts

- Strong understanding of floating-point numerical correctness and tolerance thresholds

- Experience debugging kernel compilation and runtime issues

- Ability to distinguish software defects, environment problems, and genuine optimization challenges

#### Relevant Experience

Candidates should have experience with several of the following types of work:

- Writing kernels from technical specifications

- Translating kernels between CUDA, Triton, or other frameworks

- Migrating kernels across hardware platforms

- Debugging incorrect kernel implementations

- Optimizing kernel performance

- Fusing multiple operations into optimized kernels

#### Nice to Have

- Experience across both NVIDIA GPU and custom accelerator ecosystems

- Experience with AWS Trainium, TPU, JAX, or other accelerators

- Compiler engineering experience

- Familiarity with MLIR, XLA, or intermediate representation lowering

- Contributions to GPU or ML kernel libraries

- Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls

- Experience with AI model evaluation, RLHF, or technical benchmark development

#### What You’ll Be Responsible For

- Reviewing GPU and accelerator kernel implementations for correctness

- Comparing outputs against reference implementations

- Evaluating numerical tolerance thresholds

- Reviewing kernel benchmarks and determining whether comparisons are fair

- Identifying performance bottlenecks and optimization opportunities

- Assessing whether performance targets are realistic given hardware limits

- Reviewing kernel translations and hardware migrations

- Identifying compilation, driver, memory, shape, and runtime issues

- Determining whether technical tasks are genuinely difficult or incorrectly configured

- Providing clear, actionable technical feedback

#### Engagement

**Work Type:** Remote
**Engagement:** Part-time, project-based consulting
**Focus:** GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware, **optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.**

---

Source: Anyone AI's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/anyone-ai-gpu-kernel-engineer-cuda-triton-accelerator-performance-610f982e-8adb-4fe5-9c00-d004510f3f8f
