Location: Central Singapore (Hybrid)
Employment Type: 12-month contract; renewable based on project requirements and performance
Experience: 5+ years in Production Support, Site Reliability Engineering, or DevOps within trading or financial services.
Key Stack and Practices: Trading platforms, application and infrastructure monitoring, incident and problem management, root-cause analysis, automation, release deployment, patch coordination, change management, disaster recovery, business continuity, governance, compliance, and non-financial risk controls.
Key Responsibilities:
- Provide rapid production support for trading platforms, minimising downtime and business disruption while maintaining clear stakeholder communication.
- Investigate incidents, lead root-cause analysis and blameless reviews, and implement preventive actions and automation to reduce repeat issues and manual effort.
- Maintain platform availability, resilience, and application health through proactive monitoring, alert response, and start-of-day and start-of-week operational checks.
- Configure and maintain monitoring environments and oversee critical technology processes linked to exchange obligations.
- Coordinate application releases and infrastructure patching in line with approved change controls, including occasional out-of-hours support.
- Participate in disaster recovery and business continuity exercises, including scheduled activities outside standard working hours.
- Partner with developers, project managers, and business analysts to deliver scalable, reliable, and fault-tolerant platform improvements.
- Resolve user queries and service requests within agreed service timelines.
- Follow incident, problem, governance, compliance, and operational risk requirements.
- Mentor junior team members and contribute to engineering, automation, knowledge-sharing, and process-improvement initiatives across local and global teams.