Lead Site Reliability Engineer (SRE) – DevOps & Observability
<p><span>Imagine working at Intellibus to engineer platforms that impact billions of lives around the world. With your passion and focus, we will accomplish great things together.</span></p><p><br></p><p><strong>Our Platform Engineering Team is looking for experienced DevOps / SRE leaders who can help build highly reliable, observable, and automated infrastructure supporting mission-critical applications.</strong></p><p><br></p><p><strong>We are looking for hands-on technical leaders with deep experience in Datadog, infrastructure automation, configuration management, Terraform/Chef, Bash/Shell scripting, cloud infrastructure, and production systems.</strong></p><p><br></p><p><span>The ideal candidate will also have a strong understanding of Java-based applications and Java coding, as this role will work closely with Java engineering teams and distributed application platforms.</span></p><p><br></p><p><strong>We are looking for Architects who can do the below but not limited to:</strong></p><ul><li><span>Lead DevOps/SRE initiatives across mission-critical environments.</span></li><li><span>Design and implement observability solutions using Datadog.</span></li><li><span>Build dashboards, monitors, alerts, logs, traces, and actionable operational metrics.</span></li><li><span>Define and improve SLIs, SLOs, SLAs, and error-budget practices.</span></li><li><span>Automate infrastructure provisioning and configuration using Terraform, Chef, or similar tools.</span></li><li><span>Develop and maintain Bash/Shell scripts for infrastructure and operational automation.</span></li><li><span>Manage deployment, configuration, and environment automation across development, QA, and production.</span></li><li><span>Troubleshoot complex production, infrastructure, networking, and application issues.</span></li><li><span>Improve system availability, scalability, performance, and reliability.</span></li><li><span>Support CI/CD pipelines and automated deployments.</span></li><li><span>Work closely with Java engineering teams to understand application behavior, performance, dependencies, and production issues.</span></li><li><span>Analyze Java applications from an operational perspective, including JVM behavior, memory, CPU, threads, logs, and application performance.</span></li><li><span>Participate in incident response, root-cause analysis, and post-mortems.</span></li><li><span>Identify opportunities to eliminate manual processes through automation.</span></li><li><span>Establish operational standards, runbooks, and best practices.</span></li><li><span>Provide technical leadership and mentor other DevOps/SRE engineers.</span></li></ul><p><br></p><p><strong>Core DevOps / SRE Responsibilities</strong></p><p><span>Observability</span></p><ul><li><span>Hands-on Datadog experience is required.</span></li><li><span>Build and maintain dashboards and actionable alerts.</span></li><li><span>Monitor applications, infrastructure, services, APIs, and databases.</span></li><li><span>Configure APM, logs, metrics, traces, and service-level monitoring.</span></li><li><span>Identify performance and reliability issues before they impact clients.</span></li><li><span>Infrastructure Automation</span></li><li><span>Design and maintain infrastructure using Infrastructure as Code (IaC).</span></li><li><span>Strong experience with Terraform and/or Chef.</span></li><li><span>Automate configuration management and environment provisioning.</span></li><li><span>Manage infrastructure consistency and configuration drift.</span></li><li><span>Develop reusable automation frameworks and modules.</span></li><li><span>Scripting & Automation</span></li><li><span>Strong Bash/Shell scripting experience.</span></li><li><span>Automate deployments, operational processes, monitoring, and infrastructure tasks.</span></li><li><span>Python scripting is a plus.</span></li><li><span>Ability to troubleshoot scripts and automation failures in production.</span></li><li><span>Java Application Understanding</span></li><li><span>This is not a Java Developer position, but candidates must have a strong understanding of Java-based applications.</span></li></ul><p><br></p><p><strong>Should be able to:</strong></p><ul><li><span>Read and understand Java code.</span></li><li><span>Troubleshoot Java application issues from an infrastructure/SRE perspective.</span></li><li><span>Understand JVM, memory, CPU, threads, garbage collection, and application performance.</span></li><li><span>Work effectively with Java/Spring Boot engineering teams.</span></li><li><span>Understand REST APIs, microservices, and distributed applications.</span></li><li><span>Cloud & Platform Engineering.</span></li></ul><p><br></p><p><strong>Key Skills & Qualifications:</strong></p><ul><li><span>12+ years of experience in DevOps, SRE, Infrastructure Engineering, Platform Engineering, or related roles.</span></li><li><span>8+ years of hands-on DevOps/SRE leadership experience.</span></li><li><span>Strong hands-on Datadog experience.</span></li><li><span>Strong experience with Terraform and/or Chef.</span></li><li><span>Strong Bash/Shell scripting skills.</span></li><li><span>Strong Linux/UNIX experience.</span></li><li><span>Strong cloud infrastructure experience.</span></li><li><span>Strong understanding of Java applications and Java coding.</span></li><li><span>Experience supporting distributed systems and microservices.</span></li><li><span>Experience with CI/CD and deployment automation.</span></li><li><span>Experience troubleshooting production environments.</span></li><li><span>Strong understanding of networking fundamentals.</span></li><li><span>Experience with monitoring, logging, alerting, and observability.</span></li><li><span>Experience leading technical initiatives and mentoring engineers.</span></li><li><span>Excellent communication and problem-solving skills.</span></li></ul><p><br></p><p><strong>We work closely with</strong></p><ul><li><span>Kubernetes</span></li><li><span>Docker</span></li><li><span>AWS</span></li><li><span>Jenkins</span></li><li><span>Kafka</span></li><li><span>Splunk / ELK</span></li><li><span>Prometheus / Grafana</span></li><li><span>New Relic</span></li><li><span>CloudWatch</span></li><li><span>Python</span></li><li><span>Spring Boot</span></li><li><span>Microservices</span></li><li><span>Event-driven architecture</span></li><li><span>Performance engineering</span></li><li><span>Incident management / SRE practices</span></li></ul><p><br></p><p><strong>Compensation: $75 - $85 / Hour</strong></p><p><br></p><p><strong>Our Process</strong></p><ul><li><span>Schedule a 15 min Video Call with someone from our Team</span></li><li><span>1 Proctored GQ Test (< 60 Minutes) & Google Doc Presentation</span></li><li><span>30-45 min Final & Technical Video Interview</span></li><li><span>Receive Job Offer</span></li></ul><p><br></p><p><span>If you are interested in reaching out to us, please apply, and our team will contact you within the hour.</span></p>