Djibril Faty
15 years operating critical platforms (SLAs, 24/7 on-call, P1 incidents) and a Cloud Native specialization (Kubernetes, Terraform, GitOps, SRE). Focused on DevOps transformations in demanding environments, especially Sovereign Cloud. I bring a reliability culture to teams industrializing their practices.
Cloud Engineer / Site Reliability Engineer (SRE)
My Approach
These three values guide my day-to-day work on critical production environments.
Reliability & Availability
15 years ensuring high availability and Disaster Recovery for critical production platforms: SLA compliance, 24/7 on-call, and P1 incident management, with a strong focus on service continuity.
Automation & SRE
Monitoring, logging, and automation (Ansible, scripting) to make operations more reliable, with an SRE approach to incident management.
Cloud Native & Sovereign
AWS Certified SysOps Administrator with a Sovereign Cloud / SRE specialization: Kubernetes, Terraform, and GitOps to industrialize reliable platforms, including on sovereign infrastructure (OVH, Proxmox).
Skills & Tools
The technologies and tools I master, drawn from 15 years of production operations and my Sovereign Cloud / SRE specialization.
Cloud & Containerization
Public and sovereign cloud infrastructure, virtualization, and container orchestration, down to Kubernetes networking.
AWS
OVH
Proxmox
Kubernetes
Helm
Docker
Ingress-NGINX
Traefik
cert-manager
Cilium
Infrastructure as Code & Automation
Infrastructure provisioned and configured as code: reproducible, version-controlled, and automated.
Terraform
Ansible
Python
Bash
CI/CD & GitOps
Continuous integration and delivery pipelines, GitOps deployments, and centralized management of secrets and images.
Git
GitLab CI/CD
GitHub Actions
ArgoCD
Vault
Harbor
Observability & SRE
Metrics, alerting, and centralized logging in service of high availability, with rigorous incident management through to root cause analysis.
Prometheus
Grafana
Alertmanager
Centralized logging
High availability & DR
RCA
Systems, Networks & Data
Linux administration, networking fundamentals, and distributed block and object storage for Kubernetes.
Linux (RHEL)
TCP/IP & DNS
Load balancing
SSH / SFTP
Longhorn
MinIO
SeaweedFS
My Journey
15 years of experience on critical production platforms, from engineering school to Cloud Engineering.
Sovereign Cloud Engineer / SRE Specialization
Intensive 450-hour production-oriented program: Kubernetes, Terraform, GitOps, observability, SRE practices, and Sovereign Cloud. One-month capstone project defended before a professional jury.
Cloud Certification — AWS SysOps Administrator
AWS SysOps Administrator certification, complementing hands-on Kubernetes, Terraform, and Ansible expertise.
Operation & Support Engineer
Led support for Cloud platforms (OVH, AWS), ensuring reliability, high availability, and Disaster Recovery across Big Data client accounts (geolocation, contextual marketing, SMS Gateway), with an SRE approach: monitoring, logging, automation, and incident management.
Technical Validation Lead
Integration and validation of IPTV/VOD service platforms, end-to-end interconnection testing, non-regression test automation, and coordination of internal and partner teams.
Contact
Feel free to reach out for any question or collaboration opportunity.
linkedin.com/in/dfaty
Connect on LinkedInLet's Collaborate
A question, a project, or just want to talk Cloud and system reliability? Feel free to reach out.
Contact Me