A logo

Sr Site Reliability Engineer

Artech LLCWestbrook, ME

$63 - $70 / hour

Automate your job search with Sonara.

Submit 10x as many applications with less effort than one manual application.1

Reclaim your time by letting our AI handle the grunt work of job searching.

We continuously scan millions of openings to find your top matches.

pay-wall

Overview

Remote
On-site
Compensation
$63-$70/hour

Job Description

Introduction

Senior Site Reliability Engineer to focus on the health of a cloud-native event-driven enterprise transactional system supporting a billion-dollar line of business — its reliability, observability, performance, and resilience. As an embedded member of a large feature development team, this role brings dedicated, proactive attention to keeping the system healthy and operable, so that quality and stability advance in step with new features rather than trailing behind them.

Required Skills & Qualifications

  • 7 years in an SRE, DevOps, infrastructure, or software engineering role.
  • Experience tuning application performance based on real-world inputs — memory/CPU, message queues, request load — and adjusting resources and scaling accordingly.
  • Proven, hands-on load and performance testing experience — designing realistic scenarios, executing load/stress tests, and interpreting results to drive capacity and performance decisions (Gatling preferred; comparable tools such as JMeter, k6, or Locust welcome).
  • Hands-on Kubernetes experience.
  • Strong Terraform / infrastructure-as-code experience.
  • Solid working knowledge of DataDog — comfortable enough to be the team’s go-to and coach others on it.
  • Comfortable defining and implementing SLOs and tuning alerting.
  • Strong troubleshooting across the stack (we run Spring Boot services and an Angular frontend).
  • Demonstrated success defining SRE and operational practices and driving their adoption across a development team.
  • Genuine interest in applying AI to engineering work.
  • Prior work experience at client or in client's Industry

Applicants must be able to work directly for Artech on W2.

Preferred Skills & Qualifications

  • Experience operating on GCP / GKE.
  • Familiarity with Kafka/Pub-Sub-style messaging, Hazelcast, or similar distributed-systems components.
  • Hands-on experience with AI coding agents (e.g. Claude Code).
  • Practical experience with chaos engineering concepts — fault injection, game days, and resilience testing — and the tooling that supports them (e.g. Chaos Mesh, Gremlin, LitmusChaos). Hands-on familiarity with the concepts matters more than having run a formal program.
  • Performance-engineering depth: profiling, APM-driven bottleneck analysis, and capacity modeling.

Day-to-Day Responsibilities

  • Champion and deepen our use of DataDog across Logging, Error Tracking, RUM, Incident Management, and Case Management, and coach team members to raise their fluency with it.
  • Improve signal-to-noise on alerts and errors; define and implement SLOs; make detecting and triaging issues faster.
  • Own infrastructure-as-code: author and maintain Terraform and configure our Kubernetes service operator to manage and deploy infrastructure.
  • Analyze infrastructure and application performance; tune sizing and scaling for cost-vs-performance, including refactoring application code where it improves reliability or performance.
  • Own and evolve our Gatling load-testing framework: design realistic, high-volume load and stress scenarios, run them regularly, and translate the results into capacity, scaling, and performance decisions.
  • Establish and lead our chaos engineering practice from the ground up — design and run fault-injection experiments and game days to validate resilience and systematically harden the system.
  • Look for opportunities to embed AI/Claude into tooling and workflows to reduce manual toil.
  • Monitor infrastructure cost trends and drive efficiency improvements.
  • Participate in the on-call rotation: acknowledge alerts, triage, and coordinate the right people to remediate.
  • Act as a leader during incident response — coordinating the response, ensuring stakeholders are kept informed, and pulling in the right people to resolve incidents within our recovery time objectives (RTOs).
  • Triage production alerts and Tier-3 escalations during business hours; route issues and engage team members as appropriate.
  • Document defects well in Jira (clear repro steps, recordings where useful) and initiate/coordinate post-mortems for critical incidents.
  • Participate fully in team ceremonies (planning, grooming, retros).
  • Help shape the SRE backlog — bringing your experience to bear on what we prioritize.
  • Help define observability and supportability requirements for new features.
  • Champion reliability and resilience practices and bring the rest of the team along.

For immediate consideration please click APPLY to begin the screening process with Alex.

Automate your job search with Sonara.

Submit 10x as many applications with less effort than one manual application.

pay-wall

FAQs About Sr Site Reliability Engineer Jobs at Artech LLC

What is the work location for this position at Artech LLC?
This job at Artech LLC is located in Westbrook, ME, according to the details provided by the employer. Some roles may also include multiple work locations depending on the requirement.
What pay range can candidates expect for this role at Artech LLC?
Candidates can expect a pay range of $63.33–$70 per hour for this role.
What employment applies to this position at Artech LLC?
The employer has not provided this information. This may be discussed during the hiring process.
What is the process to apply for this position at Artech LLC?
You can apply for this role at Artech LLC either through Sonara's automated application system, which helps you submit applications 10X faster with minimal effort, or by applying manually using the direct link on the job page.