4Bell Technology

Staffing & Recruiting

SRE (Site Reliability Engineer) & Platform Engineer(R-1771)

100,000.00-2,600,000.00/A

Any Degree

IT (Information Technology)

Full-time

Kolkata

17-Aug-2026

Devops Aws Kubernetes Terraform Datadog Ci/cd ArgoCD Atlantis SLO/SLI Karpenter

Job Description

Key Operational Challenges

  • Infrastructure as Code: Terraform has become unmanageable due to drift reconciliation and multi-person collaboration issues
  • Legacy Environments: Some environments set up entirely manually, never in Terraform, with major operational challenges to migrate while keeping them online
  • Terraform Migration: Proven migration patterns in test environments, but moving to production challenging due to real-time workload requirements
  • Alert Management: Receiving too many alerts, need prioritization and structured approach to reduce noise and recategorize/adjust thresholds
  • Alert Distribution: Historically all alerts went to one tech ops team instead of being distributed to five different dev teams; working to shift left closer to developers
  • Service Catalog: Still defining service catalog and ownership model
  • Runbooks: Exist in Datadog but need more polish and structure
  • Database Scaling: Operational challenges with database scaling being addressed through data store segmentation
  • Monitoring Costs: Datadog is expensive; moved from CloudWatch ~8 years ago due to cost

Manual Changes: Infrastructure changes often implemented manually first, then imported to Terraform and rolled out to other environments

Current Tech Stack & Infrastructure

  • Cloud Infrastructure: AWS (ElastiCache, EKS, RDS Aurora MySQL)
  • Orchestration: Kubernetes on EKS with Karpenter for node management
  • Infrastructure as Code: Terraform (multiple repositories, experiencing drift and collaboration challenges)
  • Monitoring & Alerting: Datadog for monitoring, alerting, incident management, and runbooks
  • CI/CD: GitHub Actions with Atlantis for infrastructure PRs, Bitbucket Pipelines for Helm deployments, Rundeck for scripted operations
  • GitOps: ArgoCD for Kubernetes workloads
  • Data Infrastructure: Transitioning to segmented data stores with ClickHouse, MySQL Aurora, and Pulsar for event streaming
  • APM: Limited Datadog APM usage due to cost ($45 per host/month, ~10 hosts)
  • Additional Monitoring: Started Prometheus clusters within EKS for more verbose metrics alongside Datadog