Cloud Platform System Engineer
Job Title: Cloud Platform System Engineer
Key Skills: Python, Containerization, Linux Administration, CI/CD Pipelines, Cloud Monitoring & Observability, Performance
Location: Brazil
Mode: Remote
We at Coforge are seeking a highly skilled and experienced Cloud Platform System Engineer to join our team. This role focuses on performance assessment, system limits evaluation, and KPI-driven analysis for a Private Cloud Platform solution designed for distributed Edge computing. You will be part of a talented team responsible for delivering a robust private cloud platform that supports containerized applications, virtual machines, and bare metal nodes.
Key Responsibilities:
Define and maintain Key Performance Indicators (KPIs) that characterize the platform's performance boundaries and system limits, including scalability ceilings, resource saturation thresholds, latency budgets, and throughput baselines.
Track KPI trends across builds and releases, identifying performance regressions, drift, or degradation patterns early in the development cycle.
Execute performance assessments as part of pre-release validation, providing data-driven evidence for release readiness decisions.
Design and develop test workloads, synthetic applications, and stress scenarios to stimulate the system under controlled conditions and expose performance characteristics.
Develop and maintain tools, frameworks, and automated pipelines for continuous performance measurement, data collection, and KPI reporting.
Execute and analyze performance experiments (load, soak, spike, scalability, endurance) to identify system limits and regression points.
Analyze performance-related issues reported from production to identify gaps in current KPIs, test workloads, or measurement coverage, and incorporate findings into the assessment process for subsequent releases.
Support reproducibility of performance results by documenting environments, configurations, workloads, and measurement methodologies.
Troubleshoot performance anomalies and regressions by correlating system metrics, logs, traces, and resource utilization data across distributed components.
Collaborate with development and test teams to provide performance insights that inform architecture decisions, capacity planning, and release readiness.
Contribute to a highly available, carrier-grade private cloud platform aimed to be at the core of 5G and distributed Edge deployments worldwide.
Required Skills & Qualifications:
Proficiency in Python for developing automation scripts, tools, or pipelines for performance data collection, analysis, or reporting.
Strong analytical, troubleshooting, and attention to detail skills, with a data-driven and methodical approach to performance investigation.
Strong Linux familiarity, including terminal usage and understanding of OS internals (process scheduling, memory management, file systems, I/O subsystems, cgroups, namespaces).
Experience with Linux performance observability tools (e.g., perf, top/htop, vmstat, iostat, sar, dstat, pidstat, strace, bpftrace).
Experience collecting, correlating, and interpreting system metrics (CPU, memory, disk I/O, network throughput/latency) in distributed architectures.
Strong understanding of computer networking at the transport (TCP/UDP) and network (IP) layers, with the ability to assess network performance characteristics (bandwidth, latency, packet loss, jitter).
Experience installing, deploying, and configuring systems and applications on Linux-based hosts.
Able to communicate in English at a minimum according to the parameters of level B2 for understanding, speaking, and writing of the CEFR matrix.
Preferred Skills:
Experience defining KPIs for infrastructure or platform systems.
Experience with Linux performance testing frameworks or load generation tools (e.g., fio, iperf, qperf, cyclictest).
Experience building CI/CD pipelines that integrate performance benchmarking (e.g., Jenkins, GitLab CI, GitHub Actions).
Kubernetes cluster administration and understanding of K8s performance dimensions.
OpenStack administration and performance considerations.
Experience with monitoring and observability stacks (e.g., Prometheus, Grafana, ELK/OpenSearch, collectd, telegraf).
Understanding of Docker/Containers and their performance implications.
Knowledge of cloud platform concepts and capacity planning.
Experience with large-scale systems monitoring and alerting.
Posted On: 21-08-2026
At Coforge, we hire professionals based solely on their skills and do not discriminate based on age, disability, religion, gender, sexual orientation, socioeconomic status, or nationality.