Resolving Persistent vSAN Performance Issues Without Hardware Replacement

The situation

An enterprise manufacturing customer was experiencing ongoing performance complaints from application owners running on a VxRail cluster. Users reported intermittent slowness and inconsistent response times, particularly during peak business hours.

From a high level, the environment appeared healthy. Hardware alerts were minimal, and standard monitoring did not point to a clear fault. Despite this, the performance issues persisted and confidence in the platform was eroding.

The challenge

vSAN performance problems are often difficult to isolate because they can stem from multiple layers that appear “healthy” when viewed independently.

In this case:

  • There were no obvious disk failures
  • Hosts were online and responsive
  • No single alert explained the end-user experience

The risk for leadership was that the issue would be misdiagnosed as a hardware limitation, leading to unnecessary expansion or refresh recommendations without addressing the real cause.

Maven’s approach

Maven approached the issue from an operational and architectural perspective rather than assuming a component failure.

Our work focused on:

  • Reviewing vSAN performance metrics over time to identify patterns related to load and contention
  • Evaluating cache and capacity tier behavior to determine whether write buffering or read caching was being saturated
  • Reviewing storage policies to confirm they aligned with the cluster’s actual capacity and workload profile
  • Identifying conditions that can trigger excessive resync activity or background operations that compete with production workloads

Rather than making immediate changes, Maven validated each finding against the customer’s risk tolerance and operational requirements before recommending corrective actions.

The outcome

  • Performance stabilized without replacing hardware or expanding the cluster
  • Latency during peak periods was reduced to acceptable levels
  • The customer gained a clearer understanding of how their workloads interacted with vSAN and how to avoid similar issues in the future

Most importantly, the remediation focused on restoring predictable performance rather than masking symptoms with additional spend.

Why this mattered

For the CTO, the issue was not whether the platform could technically stay online. It was whether it could consistently support the business without escalating costs or complexity.

By identifying and addressing the underlying operational drivers of the performance issue, Maven helped the customer extend the useful life of the existing environment and avoid unnecessary capital expense.