4Bell Technology
SRE (Site Reliability Engineer) & Platform Engineer(R-1771)
100,000.00-2,600,000.00/A
Any Degree
IT (Information Technology)
Full-time
Kolkata
17-Aug-2026
Devops Aws Kubernetes Terraform Datadog Ci/cd ArgoCD Atlantis SLO/SLI Karpenter
Job Description
Key Operational Challenges
- Infrastructure as Code: Terraform has become unmanageable due to drift reconciliation and multi-person collaboration issues
- Legacy Environments: Some environments set up entirely manually, never in Terraform, with major operational challenges to migrate while keeping them online
- Terraform Migration: Proven migration patterns in test environments, but moving to production challenging due to real-time workload requirements
- Alert Management: Receiving too many alerts, need prioritization and structured approach to reduce noise and recategorize/adjust thresholds
- Alert Distribution: Historically all alerts went to one tech ops team instead of being distributed to five different dev teams; working to shift left closer to developers
- Service Catalog: Still defining service catalog and ownership model
- Runbooks: Exist in Datadog but need more polish and structure
- Database Scaling: Operational challenges with database scaling being addressed through data store segmentation
- Monitoring Costs: Datadog is expensive; moved from CloudWatch ~8 years ago due to cost
Manual Changes: Infrastructure changes often implemented manually first, then imported to Terraform and rolled out to other environments
Current Tech Stack & Infrastructure
- Cloud Infrastructure: AWS (ElastiCache, EKS, RDS Aurora MySQL)
- Orchestration: Kubernetes on EKS with Karpenter for node management
- Infrastructure as Code: Terraform (multiple repositories, experiencing drift and collaboration challenges)
- Monitoring & Alerting: Datadog for monitoring, alerting, incident management, and runbooks
- CI/CD: GitHub Actions with Atlantis for infrastructure PRs, Bitbucket Pipelines for Helm deployments, Rundeck for scripted operations
- GitOps: ArgoCD for Kubernetes workloads
- Data Infrastructure: Transitioning to segmented data stores with ClickHouse, MySQL Aurora, and Pulsar for event streaming
- APM: Limited Datadog APM usage due to cost ($45 per host/month, ~10 hosts)
- Additional Monitoring: Started Prometheus clusters within EKS for more verbose metrics alongside Datadog