Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Infrastructure Operations Engineer

Kayak
CompanyKayak
CategoryEngineering
LocationConcord
RemoteOn-site
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted1 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
KAYAK, part of Booking Holdings (NASDAQ: BKNG), is a leading travel search engine. With billions of queries across our platforms, we help people find their perfect flight, stay, rental car and vacation package. We're also transforming business travel with a new corporate travel solution, KAYAK for Business. As an employee of KAYAK, you will be part of a travel company that operates a portfolio of global metasearch brands including momondo, Cheapflights and HotelsCombined, among others. From start-up to industry leader, innovation is in our DNA and every employee has an opportunity to make their mark. Our focus is on building the best travel search engine to make it easier for everyone to experience the world. In this role, you'll join KAYAK's Operations team in our Concord, MA office and play a key role in the day-to-day support of our development and production environments. You'll work closely with engineering, security, and platform teams to keep our infrastructure reliable, performant, and ready to scale. If you’re a team player who loves to learn, engage with new tech, and help build better- we’d love to hear from you! In this role, you will: - Receive, triage, and prioritize inbound tickets from developers and business teams, ensuring timely resolution and clear communication throughout the lifecycle of each request. - Serve as a primary point of contact for infrastructure-related incidents, driving root cause analysis (RCA) and implementing corrective actions to prevent recurrence. - Monitor, audit, and continuously improve the health and performance of our production hosting platform using tools such as LogicMonitor, Kibana, and Elasticsearch. - Proactively identify anomalies, performance bottlenecks, and capacity risks in our infrastructure, escalating and coordinating remediation with relevant engineering teams. - Maintain, test, and refine backup systems and data retention policies to ensure business continuity and compliance with internal standards. - Develop and document operational runbooks, standard operating procedures (SOPs), and post-incident reports to build institutional knowledge and improve team efficiency. - Collaborate cross-functionally with software engineering, security, and platform teams to support infrastructure changes, deployments, and release processes. - Contribute to longer-term infrastructure improvement projects — from automation initiatives to platform migrations — as a hands-on team resource. - Participate in on-call rotations to support 24/7 production environment availability, responding to critical alerts and escalations as needed. - Identify opportunities to automate repetitive operational tasks using scripting (Bash, Python, etc.) to reduce toil and improve team velocity. - Support capacity planning efforts by tracking resource utilization trends and making data-informed recommendations for scaling infrastructure. In this role, you will not: - Perform IT tasks for users like replacing mice, keyboards, etc. - Manage logins, groups, and email accounts for internal users. Please apply if you have: - A Bachelor’s degree in Computer Science, Information Systems, or a related technical field (or equivalent practical experience) - 3+ years of hands-on experience in an infrastructure, systems, or platform operations role in a production environment - Strong working knowledge of Linux systems administration (RHEL, CentOS, Ubuntu, or similar) and proficiency in shell scripting (Bash); Python scripting experience is a strong plus. - Solid understanding of datacenter operations including physical/virtual server management, networking fundamentals (DNS, TCP/IP, load balancing), and storage systems. - Demonstrable experience with monitoring and observability tooling — such as LogicMonitor, Datadog, Prometheus, or equivalent — including alerting, dashboarding, and threshold tuning. - Hands-on experience with log aggregation and analysis
HOUSE ADYour CV gets thirty seconds.CV writing and honest review. English & Greek.kaeros.app →