Required Skills: Datadog, Dynatrace, Open Telemetry, Golang, Python, Java
Job Description
Required Skill Set
Core Technical Skills- Observability Platforms
- Datadog, Dynatrace, Prometheus, Grafana, OpenSearch / Elasticsearch
- Jaeger, Tempo, Open Telemetry
Programming & Development
- Golang (Preferred), Python, Java and C#
Cloud & Platform Engineering
- Kubernetes, Cloud Platforms (AWS/Azure/GCP), Containerized Environments
- Distributed Systems Architecture
Observability Engineering
- Metrics, Logs, Traces, Events, SLI / SLO Design, Telemetry Collection & Processing
- Instrumentation of Applications, Alerting & Monitoring Frameworks and Telemetry Pipeline Development
Automation & AIOps
- Observability Automation, AIOps Solutions, API Development, Internal Tooling Development and MTTR Reduction Initiatives
Engineering Experience Required
- 3 to 5 years of hands-on Observability Engineering experience
- Strong Software Engineering background
- Production support and troubleshooting experience
- Experience operating large-scale distributed systems
- Performance tuning and optimization and Debugging production incidents using telemetry data
Nice-to-Have Skills
- Open Telemetry Collectors, Exporters, Processors, SDKs
- Site Reliability Engineering (SRE), Platform Engineering
- Internal Developer Platforms, High-volume Data Pipelines
- Streaming & Event Processing, Buffering and Backpressure Management
Key Responsibilities
- Design and build observability platforms and services
- Develop telemetry pipelines and instrumentation libraries
- Implement Open Telemetry solutions
- Improve reliability, debuggability, and operational visibility
- Define observability standards and best practices
- Optimize telemetry systems for scale, cost, and reliability
- Collaborate with Application, Infrastructure, Security, Operations, and SRE teams
Ideal Candidate Profile
✅ Strong Observability Tools Expertise (Datadog, Dynatrace, Open Telemetry)
✅ Software Development Experience (Golang/Python/Java)
✅ Platform / Cloud Engineering Background
✅ Kubernetes & Distributed Systems Knowledge
✅ Automation and AIOps Mindset
✅ Experience troubleshooting complex production environments