AI Ops Engineer
Parspec
| Company | Parspec |
| Category | Engineering |
| Location | Bengaluru |
| Remote | Hybrid |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 10 Mar 2026 |
| Last verified | 31 Jul 2026 |
| Source | Employer career page (ashby) |
Description
ABOUT PARSPEC
Parspec is building the AI and digital infrastructure for the construction materials supply chain.
Construction is a $15 trillion industry, yet the systems that underpin the buying and selling of materials remain fragmented, manual, and disconnected. Distributors and rep agencies rely on spreadsheets, PDFs, phone calls, and siloed tools to find new products and quote and manage projects; creating delays, errors, and margin erosion across the supply chain.
Parspec is an AI-native platform that powers how construction products are discovered, bought, and sold. Trusted by more than 300 MEP distributors and rep agencies, Parspec helps project-driven businesses bid faster, win more work, and operate more profitably. By combining product intelligence, AI-powered workflows, and a connected ecosystem, Parspec is laying the foundation for a more intelligent, efficient construction supply chain.
Founded in 2021 and headquartered in San Mateo, California, Parspec has raised $31 million from leading deep-tech and construction-technology investors.
THE OPPORTUNITY
We are looking for an experienced AIOps / LLMOps Engineer to help design, deploy, and manage the AI infrastructure that powers Parspec’s next-generation AI systems.
This role will focus on building and maintaining scalable, secure, and observable AI platforms that support generative AI applications and document intelligence workflows. You will work on self-hosted large language models, asynchronous inference pipelines, and production-grade ML infrastructure on AWS.
You will collaborate closely with AI researchers, product teams, and backend engineers to ensure reliable and efficient deployment of AI systems across Parspec’s platform.
Preferred location: Bengaluru, with regular in-office collaboration.
WHAT YOU WILL ACHIEVE AND KEY RESPONSIBILITIES
AI INFRASTRUCTURE & LLM PLATFORM DEVELOPMENT
- Design and build document AI platforms powered by generative AI, leveraging asynchronous architectures for scalable inference.
- Implement event-driven and queue-based systems to support elastic scaling and non-blocking AI workflows.
- Architect and maintain self-hosted LLM infrastructure using tools such as vLLM or Ollama on Kubernetes or EC2 with GPU orchestration.
LLM OPERATIONS & MODEL GOVERNANCE
- Manage production systems for LLM serving, inference pipelines, and AI workflow orchestration.
- Implement LLM gateways and routing systems (e.g., LiteLLM, Portkey) to ensure proper model usage and governance.
- Develop guardrails and monitoring systems to reduce hallucinations, misuse, and unsafe outputs in generative AI systems.
OBSERVABILITY & AI SYSTEM MONITORING
- Implement end-to-end observability for AI/ML pipelines using distributed tracing and monitoring tools.
- Monitor AI system health using platforms such as OpenTelemetry, AWS X-Ray, Prometheus, and Grafana.
- Track performance metrics including latency, token usage, inference quality, and model drift.
ML PLATFORM & WORKFLOW MANAGEMENT
- Manage machine learning workflows using tools such as MLflow, Kubeflow, or SageMaker MLFlow setups.
- Enable experiment tracking, model versioning, and deployment pipelines for production AI systems.
- Work closely with engineering teams to integrate AI workflows into scalable backend systems.
SECURITY & INFRASTRUCTURE OPTIMIZATION
- Implement AI platform security controls including Bedrock Guardrails, KMS encryption, IAM least-privilege policies, VPC endpoints, and CloudTrail auditing.
- Optimize AWS infrastructure—including Bedrock, SageMaker, and EKS—for cost efficiency, performance, and reliability.
- Ensure production AI systems maintain high availability and security standards.
WHY THIS MATTERS
Generative AI systems are only as powerful as the infrastructure that supports them. Building reliable, scalable AI platforms, especially those that serve complex document intelligence workloads, requires deep experti