Stop applying to jobs that are already dead.
Every listing verified, aged honestly, expired when filled.

All listings

Anyone AI via Ashby

AWS Trainium / NKI Kernel Expert

Level not stated Worldwide
still open verified 9h ago posted 11h ago checked 1h ago
Apply at jobs.ashbyhq.com

This is the employer's own posting, not a copy on a job board.

What we know

Is it still open?

Confirmed still open

Last checked 9h ago — checked against the employer's own applicant tracking system, which is the company answering directly.

We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.

Check this listing's status as JSON

How old is it?

Posted 11h ago

The date the source published, not the day we noticed it (2026-09-15). Last seen at its source 1h ago.

Is it remote?

Argentina - Fully Remote, Ecuador - Fully Remote, Mexico - Fully Remote, Colombia - Fully Remote

That is the location the employer filed this posting under. Quoted as written — we do not re-word the source's own location.

Who may apply?

Available worldwide

The description states no restriction of its own. This is the source's own tag.

Skills named in the ad

AWSPerformance Optimization

Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.

Carried by 1 source

The listing

Anyone AI is recruiting experienced AWS Trainium / Neuron Kernel Interface (NKI) engineers for a specialized project focused on evaluating and improving kernel development tasks for AI workloads.

We’re looking for engineers with hands-on experience building or optimizing NKI kernels on AWS Trainium or Inferentia2 hardware who understand how Trainium’s architecture differs from traditional GPU programming.

What You’ll Work On

You’ll review and evaluate technical tasks involving:

  • NKI kernel correctness and Trainium-specific development patterns

  • CUDA → NKI kernel migrations

  • Trainium performance optimization and benchmarking

  • Memory management across SBUF, PSUM, and HBM

  • Tile-based computation and DMA scheduling

  • Cross-platform numerical correctness between CUDA/Triton and NKI

  • Trainium-specific performance bottlenecks and optimization opportunities

  • Technical feedback and quality assessment of kernel implementations

The work involves determining whether implementations are not only technically correct, but also idiomatic and optimized for Trainium hardware rather than simply translated from GPU-based approaches.

What We’re Looking For

  • 2+ years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface (NKI)

  • Experience working with AWS Trainium and/or Inferentia2

  • Strong understanding of:

    • Tile-based computation

    • SBUF / PSUM / HBM memory hierarchy

    • Partition dimension constraints

    • DMA orchestration

    • Trainium-specific optimization techniques

  • Ability to evaluate CUDA → NKI migrations

  • Experience profiling and optimizing workloads on Trainium

  • Understanding of numerical differences across GPU and Trainium backends

  • Strong ability to analyze complex technical implementations and provide clear written feedback

Nice to Have

  • Experience with the AWS Neuron SDK or Neuron Compiler

  • CUDA or Triton kernel development experience

  • Knowledge of NeuronCore-v2 architecture

  • Experience with FP32, BF16, FP8, and INT8 workloads

  • Experience benchmarking workloads on Trn1 or Trn2 instances

  • Familiarity with nki.language, @nki.jit, or XLA custom calls

  • Experience with technical evaluation, AI/ML data projects, RLHF, or rubric-based assessment

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: AWS Trainium / NKI kernel engineering and technical evaluation

This is a strong fit for engineers who have worked deeply with AWS Trainium infrastructure and low-level ML kernel optimization and are interested in applying that expertise to technically challenging AI projects.

Apply at jobs.ashbyhq.com