Data Site Reliability Engineer (SRE)

Clearance Level
None
Category
IT Infrastructure and Operations
Location
Remote, Working from the USA
Key Skills For Success

CI/CD

Containerization

Structured Query Language (SQL) Development

REQ#: RQ227229
Public Trust: BI Full 6C (T4)
Requisition Type: Regular
Your Impact

Own your opportunity to support the missions that matter. From working with technologies like AI, cyber and cloud to careers in intelligence and health, we offer endless opportunities to apply your expertise to create a safer, smarter world while building new skills to propel your career forward.

Job Description

Seize your opportunity to make a personal impact supporting the Case Management Modernization (CMM) Program.  The CMM program is an initiative to support the Administrative Office of the US Courts (AO) in developing a modern cloud-based solution to support all 204+ federal courts across the United States.

GDIT is your place to make meaningful contributions to challenging projects and grow a rewarding career. The Data Site Reliability Engineer (SRE) will work as part of the CMM Data Modernization and Governance team responsible for delivering integrated data governance, engineering, data platform, reporting, analytics, and Artificial Intelligence (AI)/Machine Learning (ML) capabilities that support operational decision-making and fulfill the AO's data and analytics objectives in support of the CMM program.

The successful candidate will be responsible for providing technical leadership for the day-to-day operational support, reliability, performance, and continuous improvement of the CMM data platforms, pipelines, applications, and analytics services. This role ensures that data services remain secure, available, reliable, and aligned with established service levels, data governance standards, architecture principles, and operational procedures.

THE DATA SITE RELIABILITY ENGINEER (SRE) WILL EXECUTE THE FOLLOWING RESPONSIBILITIES

  • Provide comprehensive real-time monitoring, incident and event management, capacity planning, and operational reporting to support application deployments, maintain system health, predict demand, and align cloud operations with evolving business and security objectives.

  • Maintain and audit user roles and responsibilities in cloud environments.

  • Integrate Single Sign On (SSO), Multi-Factor Authentication (MFA) and group identity management managed through the Judiciary Enterprise Network Information Exchange (JENIE) for enforcing least privilege access.

  • Adhere to guidelines prescribed by the Government and continuously assess and improve credential management processes for all user credentials.

  • Provide Disaster Recovery (DR) and Continuity of Operations (COOP) options. This must include high-availability options, including fault-tolerant and automated failover designs.

  • Integrate DevSecOps tools and processes seamlessly with enterprise systems (Integrated Development Environments (IDEs), ticketing, monitoring, etc.) to avoid fragmentation and ensure unified security posture.

  • Provide and manage a centralized secrets management system with automated rotation, access logging, and policy enforcement to securely store, manage, and control access to sensitive information and to prevent unauthorized access and data breaches for any administrative user account.

  • Integrate security tools (example: SAST, DAST, SCA, CSPM) into pipelines for continuous assessment and remediation.

  • Implement unified, automated, continuous monitoring (24/7/365) systems and tools for security, performance, and compliance across all environments, leveraging dashboards and alerting for real-time visibility. Provide supplemental monitoring of event response activities beyond normal business hours (7a.m – 6p.m Eastern Time). Systems and tools shall capture data without including a required response to alerts.

  • Ensure automated generation and management of Software Bill of Materials (SBOM) for all deployed artifacts, supporting transparency and compliance.

  • Provide diagnostics, metrics’ gathering, and performance tuning services.

  • Provide canary release function for end-user testing to support beta testing.

  • Configure an alert mechanism so that the support teams can react in an instance of unusual behavior.

  • Implement and operate a comprehensive incident and event management process, including integration with enterprise SIEM solutions, automated alerting, escalation workflows, and root cause analysis for all critical incidents.

  • Provide engineering support to ensure prompt detection, logging, diagnosis, escalation, and resolution of incidents to restore normal service operations as quickly as possible and minimize impact.

  • Perform systems support in identifying, analyzing, and eliminating the root causes of recurring incidents to minimize continued adverse impacts and potential degradation of services.

  • Make recommendations for the improvement of Incident and Problem management consistent with industry’s best practices for the cloud.

  • Maintain knowledge base of known issues, resolutions, and best practices for operational continuity.

  • Perform automated health checks across the full stack (Operating System, Application, Database and PaaS services) at agreed levels on an agreed frequency.

  • Provide a monthly issues management report. The report shall include cloud-related incidents, any stability and performance issues, configurations issues, quantity of tickets received, and time duration to resolve tickets.

  • Develop and implement thresholds, rules, and response procedures based on product team’s recommendation.

  • Monitor resource utilization (e.g., CPU, Memory, Disk Space) for the cloud hosted Virtual Machines (VMs) and other cloud services.

  • Manage the resolution procedures for any threshold breaches for cloud resources.

  • Improves system reliability, observability, automation, scalability, and operational resilience through engineering practices.

  • Monitors, maintains, and optimizes cloud infrastructure, databases, and platform services for reliability and performance.

  • Act as FinOps Analyst and perform cost optimization.

QUALIFICATIONS

  • Education: Bachelor's degree in Computer Science, Software Engineering, or related field. (Or equivalent experience.)

  • Experience: 5+ years’ experience in IT systems engineering, systems development, systems coding, and programming.

  • Deep expertise with AWS services, including monitoring, logging, compute, storage, and networking.

  • Proficiency in Infrastructure as Code (IaC) tools like Terraform, AWS CloudFormation, or Azure Bicep.

  • Hands-on experience with monitoring and APM tools such as CloudWatch, Azure Monitor, Datadog, Prometheus, Grafana, New Relic, etc.

  • Solid understanding of incident response, change management, and ITIL-based operational support.

  • Familiarity with CI/CD toolchains and automation platforms (Jenkins, GitHub Actions, GitLab, ArgoCD).

  • Strong scripting skills (Python, PowerShell, Bash) for automation and orchestration.

  • Advanced experience in providing DevSecOps implementation using GitOps, or similar tools.

  • Experienced in developing, testing, and maintaining containerized applications.

  • Expert knowledge of source version control, build/release tools and methodologies, CI/CD pipelines and the Software Build process.

  • Experience in building and maintaining CI/CD pipelines for large enterprises that consist of a large number of complex applications.

  • Ability to be flexible and work on several different products while supporting multiple teams

  • Experience with FinOps practices, cost modeling, forecasting, and optimization tools within cloud platforms.

  • Understanding of federal compliance and security frameworks (e.g., FedRAMP, NIST, JISF Rev 5).

  • Ability to analyze logs and metrics and conduct performance tuning for cloud-based services and applications.

  • Experience working across multiple product teams to get a grasp of a product and/or programs overall state of health.

  • ITIL, AWS SysOps, or Google Professional Cloud DevOps Engineer certifications are a plus.

COMMUNICATION & ORGANIZATIONAL SKILLS

  • Excellent presentation and communication skills.

  • Consultant mindset with the ability to work with high level customer stakeholders and build excellent customer relationships.

  • Experience identifying and applying industry tools, solutions, methods best practices, and emerging technologies.

  • Strong analytical skills and problem-solving skills with the ability to formulate and communicate recommendations for improvement.

  • Experience with process design and documentation methodologies, and design and production of quality deliverables, process and use case modeling, business case development.

  • Demonstrated ability to work effectively, independently, and as part of a team.

Security Clearance Level: Must be able to pass a background check to obtain a position of Public Trust.
 

Must be a US Person (Green Card Holder, US Permanent Resident Alien, Refugee, Asylee, or US Citizen).

Location: Remote.

GDIT IS YOUR PLACE
At GDIT, the mission is our purpose, and our people are at the center of everything we do.

● Growth: AI-powered career tool that identifies career steps and learning opportunities
● Support: An internal mobility team focused on helping you achieve your career goals
● Rewards: Comprehensive benefits and wellness packages, 401K with company match, and competitive pay and paid time off
● Flexibility: Full-flex work week to own your priorities at work and at home
● Community: Award-winning culture of innovation and a military-friendly workplace

OWN YOUR OPPORTUNITY
Explore an enterprise IT career at GDIT and you’ll find endless opportunities to grow alongside colleagues who share your desire to drive operations forward.

#GDITLA

Work Requirements

Years of Experience

5 + years of related experience

* may vary based on technical training, certification(s), or degree

Certification

Travel Required

Less than 10%

Salary and Benefit Information

The likely salary range for this position is $111,155 - $150,385. This is not, however, a guarantee of compensation or salary. Rather, salary will be set based on experience, geographic location and possibly contractual requirements and could fall outside of this range.
View information about benefits and our total rewards program.

Our Identity Verification Process

As part of the hiring process, we will ask you to complete an identity verification process that leverages advanced biometrics and artificial intelligence to ensure authenticity and protect against identity fraud. You are expected to be on camera during virtual interviews. We reserve the right to take your picture to verify your identity and prevent fraud. By proceeding, you authorize the collection, processing, and use of your biometric data for identity verification and security purposes.

About Our Work

We are GDIT. A global technology and professional services company that delivers technology solutions and mission services to every major agency across the U.S. government, defense and intelligence community. Our 26,000 experts extract the power of technology to create immediate value and deliver solutions at the edge of innovation. We operate across 50+ countries worldwide, offering leading mission-ready capabilities in AI, cloud, cyber and software development.

Join our Talent Community to stay up to date on our career opportunities and events at gdit.com/tc.

Equal Opportunity Employer / Individuals with Disabilities / Protected Veterans