Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

HPC Architect

Andromeda
CompanyAndromeda
CategoryUncategorised
LocationSan Francisco
RemoteRemote
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted31 Jul 2026
Last verified1 Aug 2026
SourceEmployer career page (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
HPC ARCHITECT LOCATION: NORTH AMERICA REMOTE/SF-HYBRID · FULL-TIME About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers. We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible. Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth. Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering. The Role This role owns our provider relationships on the technical side. You're the person who decides whether a provider's cluster is good enough to join the Andromeda network, and the person who helps them get there when it isn't. That means vetting prospective providers against our quality bar and then working alongside their engineers to bring their clusters onto the network cleanly. You'll work closely with our compute procurement team to identify and qualify new providers, and you'll be the standing technical relationship with the providers we already have. When procurement finds a promising provider, you're the one who validates the claims. When a provider says their fabric is non-blocking and their nodes are burn-in tested, you're the one who verifies it by reviewing their evidence, and sometimes by running the tests yourself on their hardware. The qualification bar you'll hold providers to mostly doesn't exist yet in written form. Some of it lives informally in how our SREs evaluate clusters today; much of it hasn't been defined at all. You'd be the first person in this role, so a large part of the job is building that bar: the acceptance test suite, the quality thresholds, and the technical standards a provider must meet. Then making them rigorous enough that we can trust them and clear enough that providers can build to them. This role is the technical counterpart to our Provider Technical Program Manager, who owns delivery timelines, escalations, and incident command. You own the technical judgment: is this cluster ready, what's wrong with it, and what will it take to fix. During provider incidents, command and coordination sit with the program manager and technical response sits with our SREs. Your involvement is upstream, making sure clusters that would have caused those incidents never make it onto the network. What You’ll Do - Vet prospective compute providers: assess cluster architecture, GPU hardware, network fabric, storage, and orchestration against Andromeda's quality metrics, and make the call on whether a cluster qualifies. - Define the qualification bar itself. Build the acceptance test suite, benchmark methodology, and quality thresholds from scratch, formalizing what currently exists only as SRE tribal knowledge. Own and evolve these standards as the fleet and the market change. - Run validation hands-on where it matters: burn-in testing, fabric validation (InfiniBand/RoCE), NCCL and application-level benchmarks, storage performance testing. For routine or repeat validation, define the methodology and review results rather than executing everything yourself. - Guide providers through technical onboarding: work directly with their data-center and platform engineers to remediate gaps, tune configurations, and bring clusters up to the bar on a predictable path. - Partner with compute procurement to identify and qualify new providers. Perform technical due diligence during sourcing, and a clear read on how much remediation a candid
HOUSE ADYour CV gets thirty seconds.CV writing and honest review. English & Greek.kaeros.app →
HPC Architect — Andromeda · Job Opportunities API