Jobiglo

No results.

Senior LLM Agents Architect

NVIDIA · Yoqneam

Senior 🇬🇧 English
CUDA programming Nsight Compute Nsight Systems NVIDIA GPU architecture Performance forensics

Job description

About the role

We are building AI systems that design the next generation of hardware and software. As a Senior LLM Agents Architect you will work directly with GPU architects, verification engineers, and performance experts to create end‑to‑end agentic workflows that automate kernel optimization, architectural exploration, and developer productivity.

Key responsibilities

  • Design and implement agentic AI pipelines that generate, analyze, and optimize GPU compute kernels for peak performance on NVIDIA hardware.
  • Collaborate with GPU architects to encode memory‑hierarchy trade‑offs, occupancy tuning, and instruction‑level reasoning into autonomous agents.
  • Build performance‑forensics agents that ingest large‑scale simulation traces and Nsight profiler data to pinpoint bottlenecks and suggest mitigations.
  • Partner with hardware teams to create rapid what‑if analysis flows for micro‑architecture studies such as cache sizing and compute unit scaling.
  • Prototype and productize agentic solutions, integrate with internal services, and establish evaluation backbones using offline golden sets and online telemetry.
  • Mentor teams on agent orchestration, prompting, RAG, observability, and documentation.

Required profile

  • 8+ years of experience in applied ML/AI or large‑scale systems, with at least 2 years building production LLM‑powered applications.
  • B.Sc. in Computer Science, Electrical Engineering, or a related field.
  • Deep understanding of computer architecture, especially NVIDIA GPU architecture (SMs, warp scheduling, memory model, occupancy).
  • Hands‑on experience writing, profiling, and optimizing CUDA kernels.
  • Proven ownership of an end‑to‑end agentic system or LLM application.

Required skills

  • CUDA programming and kernel optimization
  • Nsight Compute and Nsight Systems profiling tools
  • NVIDIA GPU architecture knowledge
  • Large‑language‑model (LLM) integration and agentic AI design
  • Performance forensics and data‑driven bottleneck analysis
  • Retrieval‑augmented generation (RAG) and prompt engineering

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec NVIDIA.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in Israel.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

Apply now →

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 1 month ago

Expires 1 week from now

23 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

NVIDIA

Yoqneam