Job Description
About the Role
This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.
This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership.
What You’ll Do
- Support the deployment, operation, and reliability of production services running on Kubernetes.
- Monitor service health and investigate production incidents across distributed applications.
- Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
- Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
- Support CI/CD, GitOps-based deployments, observability, and production monitoring.
- Work within a client-directed backlog and established priorities.
Qualifications
Required Qualifications
- 5+ years of experience inSite Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.
- 3+ years of hands-on production Kubernetesexperience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting
- Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
- Splunkexperience for log aggregation, search, and production troubleshooting
- Prometheus and Grafanaexperience, specifically building alert rules and dashboards, not only using existing dashboards
- CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
- Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking
- Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment. Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.
- Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication
Additional Information
Nice to Have
- Web Components / Lit experience, to perform first-level debugging of UI-related issues
- Server-side rendering or isomorphic runtime experience
- Canary rollout / multi-version production operations
- Distributed tracing and request-context correlation
- KEDA or event-driven autoscaling
- Experience with enterprise platform integration layers
What We Offer
- Competitive salary and laptop
- Professional development and training opportunities
- Work with cutting-edge cloud and container technologies
- Flexible work arrangements and collaborative team environment
- Impact on organization-wide digital transformation initiatives