System Engineer — DevOps
CloudDrove Inc.
Project: Client Project (Ad-Tech AI Platform) · Toronto, Canada (Remote)
- Provisioned and maintained AWS ECS services end-to-end — from architecting new services for feature rollouts to fixing existing services and Load Balancer (ALB) configurations — ensuring stable, production-ready deployments.
- Managed AWS IAM security posture: created and maintained IAM roles and permission policies for services and users, enforcing least-privilege access and resolving permission-related incidents across the environment.
- Drove AWS cost optimization initiatives, identifying and eliminating inefficient resource usage across compute and storage to reduce cloud spend.
- Performed routine server patching and maintenance on EC2, RDS, and ElastiCache instances, keeping infrastructure secure, up to date, and audit-ready.
- Administered a Kubernetes cluster on Oracle Cloud Infrastructure (OCI) hosting MLOps workloads — triaging pod failures, debugging cluster-level issues, and maintaining service health and uptime.
- Managed Redis caching infrastructure on Oracle Cloud VMs, ensuring 24/7 availability through proactive monitoring and rapid resolution of caching failures and downtime.
- Acted as the DevOps point of contact for client-facing engineering support, assisting with real-time application fixes, improvements, and deployment issues to keep client releases on track.
- Built and maintained Jenkins CI/CD pipelines for a 30+ microservice platform, containerizing services with Docker and managing image builds and pushes to Amazon ECR, with Helm-based deployments to AWS ECS/EKS and Oracle OKE enabling zero-downtime releases across DEV, QA, STAGING, and PROD.