Tezza Business Solutions

Computers + 1 more

Senior Platform/ Solution Architect

Job details

Contract Type

Description

Experience & Professional Profile

  • Typically 10–15+ years of experience across solution architecture, platform architecture, cloud infrastructure, capacity planning, performance engineering or related disciplines.
  • Proven experience designing and dimensioning large-scale distributed systems and microservices platforms.
  • Strong experience with Kubernetes/OpenShift and containerised application environments.
  • Hands-on experience with cloud and on-premises infrastructure architecture.
  • Demonstrable experience in capacity planning, workload modelling, performance engineering and infrastructure forecasting.
  • Experience with large-scale, high-availability, transaction-intensive environments is highly desirable.
  • Experience in telecoms, financial services, digital platforms or other high-volume technology environments is advantageous.


Core Responsibilities

  • Lead capacity planning and infrastructure dimensioning for applications, platforms and microservices-based services.
  • Translate business growth, transaction volumes and traffic forecasts into infrastructure capacity requirements.
  • Develop quantitative workload models covering normal, peak, burst and exceptional traffic conditions.
  • Determine appropriate CPU, memory, pod/replica, node and cluster requirements for application services.
  • Develop capacity forecasts and infrastructure roadmaps covering short-, medium- and long-term demand.
  • Ensure capacity plans support availability, resilience, disaster recovery and business continuity requirements.
  • Provide architecture and capacity recommendations for both cloud and on-premises environments.
  • Review existing environments to identify over-provisioning, under-provisioning, bottlenecks and capacity risks.

Microservices Capacity Planning & Dimensioning

  • Assess resource consumption and performance characteristics of individual microservices.
  • Determine minimum, normal and maximum pod/replica requirements based on workload and service-level objectives.
  • Define CPU and memory requests and limits for containers.
  • Assess horizontal and vertical scaling requirements and define appropriate scaling policies.
  • Determine node density, resource utilisation and cluster capacity requirements.
  • Account for service-to-service communication, platform overhead and infrastructure reserve capacity.
  • Establish repeatable sizing methodologies for new applications and services.
  • Validate sizing assumptions through performance and capacity testing.

Capacity Planning Parameters & Metrics

  • Define and maintain standard parameters for application and infrastructure capacity planning.
  • Analyse requests per second (RPS), transactions per second (TPS), concurrent users, sessions and transaction volumes.
  • Analyse average, peak and burst traffic and associated growth patterns.
  • Assess CPU utilisation, CPU consumption per transaction, memory utilisation, memory peaks and application heap requirements.
  • Assess pod counts, replica requirements, scaling thresholds and scaling response times.
  • Determine node CPU, node memory and allocatable cluster capacity.
  • Assess database TPS, connections, CPU, memory, IOPS and throughput.
  • Assess storage capacity, IOPS, throughput and growth.
  • Assess network bandwidth, latency and packet rates.
  • Factor in high availability, N+1/N+2 resilience, disaster recovery, growth headroom and operational reserve.

Performance Engineering

  • Lead performance engineering and capacity validation for critical applications and platforms.
  • Define and oversee load, stress, endurance, spike, scalability and capacity testing.
  • Analyse throughput, response time, latency, concurrency and resource utilisation.
  • Identify application, platform, database, storage and network bottlenecks.
  • Establish performance baselines and capacity thresholds.
  • Use performance test results to validate CPU, memory, pod, node and cluster sizing.
  • Work with engineering teams to optimise resource consumption and application performance.

Observability & Data-Driven Capacity Planning

  • Use production telemetry and historical performance data to develop evidence-based capacity models.
  • Leverage metrics, logs, traces and APM data to understand workload behaviour.
  • Use monitoring and observability platforms such as Prometheus, Grafana, OpenTelemetry, Dynatrace, AppDynamics or equivalent tools.
  • Correlate traffic, application performance, pod utilisation, infrastructure consumption and database performance.
  • Establish capacity thresholds, early-warning indicators and capacity risk dashboards.
  • Use trend analysis and forecasting to identify future infrastructure requirements before capacity constraints occur.

Architecture Governance & Standards

  • Establish standard capacity planning and dimensioning methodologies across the organisation.
  • Define architecture principles, sizing standards, resource profiles and capacity governance processes.
  • Review and approve application capacity models and infrastructure sizing proposals.
  • Ensure new services meet defined scalability, availability, performance and capacity requirements before production deployment.
  • Establish governance for capacity reviews following major releases, traffic changes or architectural changes.
  • Maintain architecture documentation, capacity assumptions, sizing models and decision records.


Start hiring with Fuzu

Recruit better talent faster - on your own or with our support.

Explore recruitment platform

Don’t miss your chance to work at Tezza Business Solutions. Enter your email to start your application now