Required Skills: SRE, DevOps, Cloud, Shell, Python,
Job Description
Technology Operations - Tech Ops - Basic
Position Overview
We are looking for an experienced Production/Application Support Engineer to join our global technology support team. The ideal candidate will be responsible for ensuring the availability, stability, performance, and reliability of production applications while providing effective incident resolution and operational support.
The role requires strong troubleshooting and analytical skills, experience with incident and change management processes, and the ability to work closely with Development, SRE, Infrastructure, and Business teams in a fast-paced production environment.
Key Responsibilities
- Provide 24/5 global production application support and ensure timely resolution of production issues.
- Monitor application health, availability, performance, and overall production stability.
- Manage Incidents, Problems, and Changes in accordance with established IT service management processes.
- Perform detailed Root Cause Analysis (RCA) for recurring and high-impact production incidents.
- Troubleshoot application, infrastructure, and integration-related issues and drive issues through to resolution.
- Support production releases, deployments, and post-release validation activities.
- Track and maintain SLA/SLO performance, including reporting and escalation of potential breaches.
- Communicate effectively with stakeholders during major incidents, service disruptions, and escalations.
- Collaborate closely with Development, SRE, Infrastructure, QA, and Business teams to resolve complex production issues.
- Identify opportunities for automation, operational improvements, and reduction of recurring incidents.
- Contribute to operational risk management, process improvements, and support documentation.
- Participate in incident reviews, problem management activities, and continuous service improvement initiatives.
- Ensure production support processes, runbooks, knowledge articles, and operational documentation are kept current.
Required Skills & Experience
- Proven experience in Application Support, Production Support, or Production Operations.
- Strong understanding of Incident, Problem, and Change Management processes.
- Hands-on experience with ticketing/service management tools such as ServiceNow, JIRA, Remedy, or similar platforms.
- Experience with application monitoring and observability tools.
- Strong troubleshooting, analytical, and problem-solving capabilities.
- Experience performing Root Cause Analysis (RCA) and resolving complex production issues.
- Exposure to Linux/Unix environments and basic command-line troubleshooting.
- Experience supporting production releases and deployments.
- Ability to work effectively in a 24/5 global support environment.
- Strong communication skills with the ability to interact with both technical and business stakeholders.
- Ability to manage multiple priorities and work effectively under pressure during critical incidents.
Preferred / Nice-to-Have Experience
- Experience within Financial Services, Banking, Trading, Investment Management, or Wealth Management.
- Knowledge of financial markets, trading platforms, or financial applications.
- Experience working with SRE, DevOps, Cloud, or Infrastructure teams.
- Familiarity with modern monitoring and observability platforms.
- Experience with automation or scripting using Shell, Python, or similar technologies.
- Understanding of high-availability, disaster recovery, and production resilience concepts.
- Experience working in a global, follow-the-sun support model.
Key Competencies
- Production Support & Service Delivery
- Incident & Problem Management
- Change & Release Management
- Root Cause Analysis & Troubleshooting
- Application Monitoring & Observability
- SLA/SLO Management
- Stakeholder & Escalation Management
- Operational Risk Management
- Continuous Improvement
- Cross-functional Collaboration
-