Lead Engineer - R01571349
- Onsite
- All software engineering jobs
- Employee
- Infrastructure Engineering
About the role
Lead Engineer
Job requirements
- Position: Site Reliability Engineer [High Performance Computing] Location: Bengaluru, India Employment Type: Full Time Consultant (FTC)Role Overview: High Performance Computing Cloud SRE Engineer The High-Performance Computing SRE squad are responsible for designing, provisioning and supporting infrastructure for large scale distributed computing (HPC Grids) in the Public Cloud. We are looking for an experienced, enthusiastic, Linux cloud infrastructure SRE engineer to join our ever-expanding team. Role responsibilities extend to:
- Migrating complex, multi-tier applications from on-premises to AWS/Azure/GSP at exceptional scale.
- Participating in operational support activities of a globally distributed team.
- Designing and operating innovative monitoring and alerting solutions.
- Troubleshooting problems and providing in-depth root cause analysis, and mitigation.
- Collaborating with Tech Risk to promote security compliance.
- Part of cloud Agile fleet.
- Working closely with service stakeholders, vendors, internal customers. Required Skills
- 5+ years of experience working with AWS and/or Azure and/or GSP.
- Advanced Kubernetes skills.
- Software Installation, configuration and patching.
- Experience with Ops support, including incident and problem management.
- Proven track record of building and supporting complex cloud infrastructure programmatically.
- Expertise in Linux OS internals and administration.
- Scripting skills in Python and/or Linux Shell.
- Infrastructure as Code tooling (Terraform, Ansible, others).
- Observability tooling and concepts (metrics, logs, alerts)
- Experience with Agile and DevOps concepts. Desired Skills
- Experience in the financial industry
- Knowledge about HPC, clustering, distributed computing.
Description as published by Brillio.