Stop applying to jobs that are already dead.
Every listing verified, aged honestly, expired when filled.

All listings

Boundless Networks, Inc. via Workable

Applied AI/ML Engineer

Level not stated $175k–250k/yr United States
still open verified 15h ago posted 43d ago checked 3h ago
Apply at apply.workable.com

This is the employer's own posting, not a copy on a job board.

What we know

Is it still open?

Confirmed still open

Last checked 15h ago — checked against the employer's own applicant tracking system, which is the company answering directly.

We re-read the employer's own applicant tracking system and the posting was still there. That is the company answering directly.

Check this listing's status as JSON

How old is it?

Posted 43d ago

The date the source published, not the day we noticed it (2026-08-03). Last seen at its source 3h ago.

Is it remote?

Marked remote on the employer's board

Their board carries a remote setting on this posting — a field they filled in, not wording we read. The location field names somewhere specific, which is usually where the team or the entity sits.

Who may apply?

United States

The description states no restriction of its own. This is the source's own tag.

Pay

$175k–250k/yr

Read out of the job description by us, not from a structured field. Shown in the posting's own currency and period; we never convert.

Skills named in the ad

GitKubernetesLLMPyTorchPython

Recognised terms only, from a fixed vocabulary — this is what CV matching compares against.

Carried by 1 source

The listing

Boundless is coordinating GPU compute at scale and building toward becoming a leader in AI. As an Applied AI/ML Engineer, you'll ship AI-powered products end-to-end on top of our growing GPU inference fleet — owning everything from serving low-latency inference to standing up reinforcement-learning post-training pipelines. This is a builder's role: you take an idea from prototype to production, tune it for throughput and cost on real GPUs, and iterate fast on customer and internal feedback.

You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.

What You'll Do

End-to-End AI Product Delivery: Own AI features and products from prototype through production — model selection, serving, evaluation, and iteration — shipping working software rather than research artifacts.

Inference Serving: Deploy and optimize LLM inference across the fleet using vLLM and SGLang. Tune continuous batching, KV-cache management, quantization, speculative decoding, and multi-model routing to maximize throughput and minimize latency and cost per token.

RL & Post-Training Harnesses: Build and operate reinforcement-learning and post-training pipelines using slime (Megatron-LM + SGLang) and Prime Intellect (prime-rl + the Environments Hub / verifiers). This includes reward and verifier design, rollout orchestration, weight synchronization, and keeping long-running training stable.

Evaluation & Iteration: Build eval harnesses and benchmarks that measure quality, throughput, and cost together, and use them to drive fast, data-informed iteration.

Work Across the Stack: Partner with Infrastructure on GPU scheduling and fleet utilization, and with Product on what to build next and why.

Requirements

  • 3+ years shipping ML/AI systems to production
  • Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM
  • Experience with RL / post-training methods (GRPO, PPO, DPO, or SFT), or strong adjacent experience and a clear desire to go deep here
  • Strong Python and PyTorch
  • Working understanding of GPU execution: batching, memory, and basic CUDA concepts
  • Comfort operating in ambiguity with a strong bias for action

Nice to Have

  • Direct experience with slime, prime-rl, the verifiers library, or Megatron-LM
  • Distributed training experience (FSDP, TP/PP/DP parallelism)
  • Quantization (FP8/INT8), P/D disaggregation, or speculative decoding
  • Experience with verifiable inference or large-scale distributed systems
  • Kubernetes and container-based deployment
  • Familiarity with GPU fleet orchestration (Ray, SkyPilot, Slurm)

Additional Requirements

  • Candidates must include a public GitHub profile in their application.
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history.
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.

Benefits

At Boundless, we take care of our people, because building the future of AI compute starts with an empowered team. Here's what you can expect when you join us:

  • Competitive salary (proposed band b/t US$175k and $250k annually) + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment

We are a global team, and applicants from around the world are welcome to apply.

Apply at apply.workable.com