Here are sample job postings for Site Reliability Engineer roles:
Software Engineer III, Site Reliability Engineering, Google Cloud
United States•Hybrid remote
About the job
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.
Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale which are unique to Google Cloud, while using your expertise in coding, algorithms, complexity analysis and large-scale system design. SRE's culture of diversity, intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.
With your technical expertise you will manage project priorities, deadlines, and deliverables. You will design, develop, test, deploy, maintain, and enhance software solutions.
Responsibilities
- Write product or system development code.
- Review code developed by other engineers and provide feedback to ensure best practices (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).
- Contribute to existing documentation or educational content and adapt content based on product/program updates and user feedback.
- Triage product or system issues and debug/track/resolve by analyzing the sources of issues and the impact on hardware, network, or service operations and quality.
- Participate in, or lead design reviews with peers and stakeholders to decide amongst available technologies.
Lead Senior Cloud Platform Engineer
Stanley Martin Homes, LLC
The Senior Cloud Platform Engineer leads the modernization of our cloud environment by driving the implementation of advanced Azure PaaS and application platform capabilities that enable the delivery of scalable, secure, and modern cloud-native applications. The Senior Cloud Platform Engineer is responsible for building scalable, enterprise-ready environments that empower our AI, Data, and Application teams to deploy workloads independently, replacing manual processes with automated, self-service APIs and CI/CD blueprints.
By abstracting the underlying complexities of Azure PaaS and Kubernetes, this position will enable a streamlined, high-performance experience for our engineering teams. The Senior Cloud Platform Engineer will build and operate these platforms, ensuring our application teams can deploy, run, and scale services reliably and securely.
Responsibilities and Duties
Platform Architecture & Modernization: Design, build, and evolve our cloud platform to provide standardized, self-service capabilities that support the modernization of our application portfolio.
Container & PaaS Orchestration: Lead the design, implementation, and operation of Azure Container Apps and Kubernetes (AKS) environments, ensuring scalable, high-performance runtime environments.
Data & AI Ecosystem Integration: Architect and automate the deployment of data-intensive platforms, including Azure Databricks and Snowflake, ensuring seamless integration with our application stack.
Engineering Enablement & Support: Work directly with development teams to debug deployment and runtime issues, improve application reliability, and drive platform stability through continuous improvements.
Incident Management: Participate in on-call rotations, incident response, and root cause analysis to maintain high platform availability and performance.
Identity & Security Architecture: Implement secure design patterns using Managed Identities and Azure Key Vault to ensure secret management and service-to-service authentication.
API-Driven Automation: Build and maintain API-driven integrations to automate provisioning and lifecycle management for application, AI, and data workloads.
Infrastructure-as-Code (IaC): Define and enforce automated standards using IaC (Terraform/Bicep) and CI/CD policy-as-code to ensure security and compliance at the point of deployment.
Application Delivery & CI/CD: Build and maintain CI/CD pipelines using GitHub Actions or Azure DevOps to standardize application deployment, environment promotion, and release automation.
Observability & Monitoring: Implement logging, monitoring, and alerting using Application Insights and Log Analytics to ensure platform health and support troubleshooting across distributed systems.
Position Requirements
- 8–12+ years in cloud engineering with a focus on application platforms, PaaS services, and container orchestration.
- Deep experience with AKS / Azure Container Apps, API Management, Service Bus / Event Grid, and Databricks / Snowflake integrations.
- Hands-on experience logging, monitoring, and diagnosing distributed systems to ensure platform health.
- Strong coding ability in Python, Go, or .NET capable of building custom tools, automation, and platform integrations.
- Solid experience with Docker and Kubernetes operations, including deployments, scaling, and deep-level troubleshooting.
- Proficient in implementing Managed Identity, Key Vault, and modern application authentication patterns.
- Expert-level skill in Terraform or Bicep, including the development of reusable, version-controlled modules.
- Proven experience building and maintaining pipelines, with a focus on integrating automated testing and security controls.
Cloud Architect
Team CARFAX is looking for hands-on, system-level thinkers who understand the many tradeoffs of technology choices, engage with distributed teams and help architect and design cloud-based platforms and data migration strategies to be leveraged across many departmental development teams.
The Tech Culture at CARFAX:
Having a creative and innovative environment where our developers can collaborate, learn and grow is something CARFAX is passionate about. We have an entire floor dedicated to our techies, designed specifically to enable teams to dream big and produce the best. Along with creating and maintaining awesome software you’ll also be able to participate in our annual Hack-a-thon or take a break by kicking back and playing the latest game on x-box when you need to re-boot the mind. Oh, and do you happen to have a dog? CARFAX is dog-friendly and no day goes by where you don’t have the chance to visit with one of the visiting pups. We even provide the dog beds, bowls and of course, toys!
At CARFAX, we believe in the power of teamwork and value in-person interactions so that we can collaborate and thrive together. This position will require 3 days per week in our Columbia, MO office subject to change with future business needs.
What you’ll be doing:
- Providing technical leadership to the data ingest development teams
- Analyze system requirements and workflow analysis, prepare design and architectural specifications
- Assist in the architectural design of migrating from legacy systems
- Ensure uniform enterprise-wide application design standards are maintained
- Participate in design decisions, including new technology research and prototyping
- Collaborate closely with other AWS engineers and architects, cloud engineers, support teams and other stakeholders
- Innovate new ideas to evolve our applications, data storage and processes
- Continuously analyze and evaluate our systems, products and process for potential improvements
- Communicate across development teams to understand pain points and places where improvements can be beneficial
- Educate business stakeholders and product development teams on the proper use of AWS products and services
You’ll need to:
- Communicate - Be vocal! We believe in the wisdom of crowds and your input is needed and valued!
- Have a Self-Driven attitude - ability to quickly come up to speed with different technologies on an ongoing basis
- Exhibit technical acumen - Be flexible! We love to change it up by using different technologies and need you to be open to new and different technologies!
- Love to learn! To get the greatest solutions we need to continually explore what’s new and be willing to dive in and learn.
What we’re looking for:
- 8+ years of professional software engineering experience with 3 years of that designing & architecting web applications
- Experience in architecting and migrating big data to AWS cost effectively, securely and reliably
- Knowledge or hands-on experience in MongoDB, Hadoop and AWS databases – RDS/Aurora, DocumentDB and DynamoDB
- Expertise in security best practices for data at rest and in transit in AWS, cloud security best practices and a DevSecOps mindset.
- Hands-on experience with AWS product offerings (e.g. EC2, VPCs, Lambda, API Gateway, etc.) and IaC tools such as Terraform or CloudFormation
- AWS Solution Architect – Professional certification is strongly preferred
- Hands-on experience migrating legacy, non-cloud applications and date to AWS including AWS data migration tooling, products and services.
- Hands-on experience building and supporting CI/CD pipelines leveraging a tool such as Jenkins or GitLab - be a firm believer in "automate everything"
- Knowledgeable of ETL and ELT architecture and patterns and best practices for large-scale performance.
- Knowledge of AWS pricing and cost estimations
- Exceptional analytical and problem-solving skills and Excellent interpersonal, collaboration and communication skills.
- Experience building web applications across diverse technologies (Java, Javascript, golang, etc.)
- Understand and advise teams in the design of common software and data architectural patterns
- Ability to facilitate technical discussions, PoC technical solutions and plan a technology roadmap
- Extensive experience in aligning application development with business needs
- Have a broad set of technology skills ingrained into DevSecOps
- Enthusiastic and motivated to apply AI to development and operations
- Operational experience - ready to own the uptime and quality of systems
- Documentation and teaching mindset - Help others!
Senior Principal Site Reliability Engineer
About the team
The SRE team at Zillow Group empowers ZG Product Teams to efficiently run “Zillow 2.0” services by reducing human error, aggressively focusing on automation, and providing deep insight into application behavior and health! We do that by incorporating aspects of software engineering and applying them to infrastructure and operations problems as a way to create and manage scalable and reliable distributed software systems.
About the role
We are looking for a Principal SRE or DevOps engineer with a demonstrated track record of building secure, large scale, highly available services using automation and Infrastructure as Code, who is well versed in cloud architecture (with a focus on Kubernetes), and loves to delight the engineers they support
As a Senior Principal SRE, you will:
- Architect, develop and deploy systems, processes and environments that support hundreds of services and microservices developed by engineers across ZG.
- Interact with the engineers we support and other internal stakeholders as a consultant and spokesperson for the work, direction and philosophy of SRE.
- Ensure developers across Zillow Group create the systems and infrastructure that powers “Zillow 2.0”.
- Consulte on design and implementation decisions, and assisting them in troubleshooting and debugging related problems.
This role has been categorized as a Remote position. “Remote” employees do not have a permanent corporate office workplace and, instead, work from a physical location of their choice which must be identified to the Company. Employees may live in any of the 50 US States, with limited exceptions. In certain cases, an employee in a remote-designated job may need to live in a specific region or time zone to support customers or clients as part of their role.
In Colorado, Connecticut, Nevada and New York City the standard base pay range for this role is $215,600.00 - $344,400.00 Annually. This base pay range is specific to Colorado, Connecticut, Nevada and New York City and may not be applicable to other locations.
In addition to a competitive base salary this position is also eligible for equity awards based on factors such as experience, performance and location. Actual amounts will vary depending on experience, performance and location.
Who you are
- 8+ years of relevant SRE, DevOps, Systems Engineering or Infrastructure Engineering experience.
- A proven track record as a technical leader with broad multi-functional impact. Previous experience as a people manager is a large plus.
- Experience with SDLC principles, architecture and operations.
- Experience with Infrastructure as Code tools and processes.
- Experience scripting/coding with Python, Java and/or Go.
- Experience writing comprehensive documentation.
- Experience working with senior leadership both inside and outside of engineering.
- Excellent written and verbal communication skills.
- Passionate about building and fostering good engineering practices and processes.