Main Responsibilities:
- Designing maintaining DevOps systems that support multiple data centers public clouds.
- Design tools mechanisms to monitmaintain cloud platform reliability, cost, performance.
- Quickly resolve online incidents through efficient SRE tools mechanisms.
- Continuously optimize ticket handling, identify high ROI manual operations that should be automated.
- Support the data science team to ensure the smooth development, experimentation, deployment, maintenance of GenAI machine learning models.
- Stay updated with the latest DevOps tools practices in Europe North America
Job Requirements:
- 3~7 years of experience in cloud operations DevOps.
- Fluent in English, capable of verbal communication meetings with European North American business teams local engineers in Vancouver.
- Familiar with CI/CD processes, with experience in designing implementing automation processes.
- Capable of monitoring maintaining cloud platform reliability, cost, security, performance.
- Understdata center localization, data isolation management, cross-border access control.
- (Bonus) Familiar with MLOps, with a certain understanding of the data science machine learning lifecycle.
岗位职责:
1. 设计和维护支持多个数据中心及公共云的DevOps系统。
2. 设计监控和维护工具,保证云平台的可靠性、成本控制及性能优化。
3. 快速响应并解决在线运行中的问题,确保业务连续性。
4. 持续优化工作流程,包括但不限于自动化操作以提高效率。
5. 支持数据科学团队,确保GenAI和机器学习模型的顺畅开发、部署及维护。
6. 保持对*新DevOps工具和实践的了解,适应不断变化的技术需求。
任职要求:
1. 具备3~7年相关工作经验,熟悉CI/CD流程及自动化工具。
2. 能够有效监控和管理云平台的运行状态,包括可靠性、安全性和性能。
3. 熟悉数据中心操作,包括数据隔离和跨国访问控制。
4. 具备良好的问题解决能力,能迅速定位并解决在线运营中的技术难题。
5. 能够与不同背景的团队成员协作,推动技术创新和优化。
6. 对数据科学有基本的理解,能够支持相关团队进行模型开发和部署。