Homelab: self-hosted infrastructure as code
75 services in production, DNS HA, full observability and real business automation, all declared as code.
75 services · HA DNS · GFS backup
- Docker Compose
- Traefik
- Prometheus
- Grafana
~ $ whoami
Site Reliability Engineer (SRE) · PagBank
availableI run mission-critical production in the financial market: incident response, observability and automation.
At home, the same engineering continues: self-hosted homelab 100% as code, with high availability and full observability.
SRE at PagBank, I sustain the resilience of an environment with 130+ microservices in production, under continuous PCI-DSS audit, with canary pipelines and automation that keep the same problem from happening twice.
I operate infrastructure end to end (network, high availability, observability, security), all declared as code, from design to root-cause analysis.
Outside the bank, the lab is real: a homelab with 75 services in production and wa-courier, the open source WhatsApp gateway I built and maintain. In parallel, I’m building a multi-tenant SaaS for the real-estate market, currently at MVP stage.
Nothing configured by hand that can’t be rebuilt from scratch straight from the repo.
Metrics, logs and alerts are part of the service design, not an extra bolted on later.
Multiple layers of defense (network, authentication, log analysis) instead of relying on a single control.
75 services in production, DNS HA, full observability and real business automation, all declared as code.
75 services · HA DNS · GFS backup
Open source HTTP gateway for WhatsApp: one number, one Docker container, zero external service. In production, MIT licensed.
MIT · multi-arch · Prometheus
Multi-tenant product in development, from the ground up. Details once the MVP matures.
MVP in progress
Senior Site Reliability Engineer
Canary deploy + tuned ASG for critical traffic. The technical gate before production.
Site Reliability Engineer · Mid-level
Self-service multi-tenant SFTP platform (Jenkins + Terraform). Zero infra queue.
Site Reliability Engineer
Started as SRE, combining troubleshooting and on-call for payment systems.
Systems Administrator · TEF
Kept Sitef, AGK and Intellinac running, automating with Salt/Ansible and Terraform.
Systems Administrator
Ran payment microservices on Docker (Mesos/Marathon), configured via Puppet.
Support Analyst
Operated and kept the on-prem infra resilient: backup, RAID, firewall, proxy (Squid), VPN and virtualization (VMware), with Shell Script automation and monitoring via Nagios. The base of SRE discipline.
Email has the best response SLA. The others are backup.
email[email protected]
githubgithub.com/thales-machado
linkedinin/thales-machado
locationPassos, Brazil · GMT-3