A SQL Server high availability discussion can turn into a technology discussion very quickly. Someone suggests an Availability Group, someone else prefers a Failover Cluster Instance, and before long the conversation is about synchronous replicas, shared storage, listeners and automatic failover.
Those may all become relevant. But at that point, one important question is often still unanswered: what exactly are we trying to protect the service against?
A failed Windows node is not the same problem as losing a storage system. An unavailable SQL Server instance is different from losing an entire site. Planned maintenance creates another requirement again. If we do not separate those situations first, choosing the technology becomes partly guesswork. We are already discussing the architecture before we have defined what it actually needs to do.
SQL Server high availability starts with the failure the service needs to survive
High availability is not a single condition. A production application can become unavailable because the server hosting SQL Server fails, because Windows or the SQL Server service has a problem, or because storage becomes inaccessible. The network path between the application and database can disappear as well. SQL Server itself may be completely healthy while clients can no longer reach it.
The scale of the required protection also varies. One environment may only need to tolerate the loss of a single physical or virtual server. Storage, networking and the data center may already provide sufficient redundancy. Another workload may need to continue operating after the loss of a complete storage platform or location. Both are legitimate requirements, but they lead to different architectures.
This is where technologies such as Failover Cluster Instances and Availability Groups enter the discussion, although they solve different parts of the problem. An FCI provides redundancy for a SQL Server instance across Windows Server Failover Cluster nodes, while the database files remain on storage accessible to those nodes. An Availability Group maintains database replicas on separate SQL Server instances and can move the primary role between them.
That difference matters, but it still does not tell us which design is appropriate. If the relevant failure is the loss of a Windows node while the shared infrastructure remains available, one architecture may be sufficient. If the requirement also includes a separate copy of the databases or protection against failure of the original storage platform, the design has to cover a wider failure domain.
The same reasoning applies beyond SQL Server. Two database servers do not automatically remove every single point of failure. Both systems may still depend on the same storage platform, network components, DNS infrastructure, authentication services or physical location. That does not make such an architecture wrong; shared components may themselves be highly available. What matters is knowing where the common dependencies are and whether they sit inside or outside the failure scenario the design needs to handle.
When reviewing an HA design, I therefore find a different question more useful than simply asking whether two SQL Servers exist: which failure could remove both systems, or make both unreachable, at the same time?
That usually gives the architecture discussion a much clearer boundary.
RPO and RTO need to be requirements, not consequences
Once we understand the failure scenarios, the next question is what the organization expects to happen after one of them occurs. How much committed data can the business afford to lose, and how long may the service remain unavailable?
We normally describe these requirements as recovery point objective and recovery time objective. They should influence the architecture before the team selects a technology. They should not become numbers derived afterwards from whatever solution the team happened to implement.
A workload that can tolerate several minutes of potential data loss gives the architecture more options. A transactional system where losing acknowledged transactions would create a serious business problem has very different requirements. The same applies to recovery time. If an application can remain unavailable for an hour, the design does not need to behave like a system expected to resume operation within a few minutes.
Availability Groups provide a good example because the configuration itself changes the recovery characteristics. Synchronous and asynchronous commit modes make different trade-offs around transaction protection, latency and distance. A synchronized synchronous replica can support failover scenarios without losing committed transactions under the conditions covered by that design. An asynchronous replica, by contrast, may not yet contain the latest transactions when the primary fails.
Microsoft also provides guidance for monitoring and estimating RPO and RTO for Availability Groups. That is useful because these values depend on the actual configuration and data movement rather than on the product name alone.
RTO depends on much more than the SQL Server role change. After a failover, the new primary must become usable, clients need to reach the correct endpoint and application connections have to recover. Listener configuration, DNS, network paths, authentication, connection strings and client retry behavior can all influence what users actually experience.
This is why I would not define recovery time as the number of seconds required for SQL Server to complete a failover. The relevant measurement is the time until the service the business depends on works again. We looked at the same distinction from another direction in SQL Server Restore Validation: a database being ONLINE is an important technical state, but it does not prove that users can actually use the application service again.
The same separation is useful for backups. High availability can reduce interruption caused by certain infrastructure failures, but it does not provide the recovery options of a properly designed backup strategy. An accidental DELETE, an unwanted application change or another logical problem can reach additional replicas as part of normal data movement. If the requirement is to return to an earlier state, that is a recovery problem, not simply an HA problem.
This distinction is also important when looking at SQL Server backup strategy. High availability and backup can both support service continuity, but they protect against different failure scenarios. Teams should not treat them as substitutes for each other.
Maintenance is part of availability too
Unplanned failures naturally dominate high availability discussions, yet many production systems experience planned maintenance far more often than catastrophic infrastructure failure. Windows needs patching, SQL Server needs updates, virtualization hosts need maintenance and hardware eventually gets replaced.
Those activities belong in the architecture discussion because the chosen design determines how the workload can move and how much disruption routine maintenance creates.
If a service may only be unavailable for a few minutes, the team needs a practical and repeatable way to move production away from the component that needs maintenance. That includes more than issuing a failover command. Someone needs to know whether the target system is ready and whether application connections move as expected. The team also needs to understand how jobs and supporting services behave and what happens if it cannot complete the planned work.
A different environment may have a generous maintenance window and no business reason to eliminate every planned interruption. In that situation, additional HA components can create complexity without solving a requirement that actually exists.
That trade-off is easy to miss when the discussion starts with technology. More redundancy can remove failure points, but it also introduces additional components, configuration and operating procedures. Availability Groups, Windows clustering, listeners, endpoints, multiple SQL Server instances and possibly multiple locations all become part of the operating model. The team needs to monitor and understand them. During normal operation that complexity may remain almost invisible, but during an incident or maintenance window it suddenly becomes part of the recovery process.
For that reason, I consider operational capability part of the architecture rather than something to document afterwards.
This becomes particularly important in environments that have been running for many years. People change roles, applications evolve, teams replace infrastructure and the original reasons behind technical decisions become less obvious. An Availability Group or cluster may still be healthy according to monitoring while nobody is entirely comfortable performing a controlled failover because the team no longer fully understands the dependencies around it.
The technology has not necessarily become unreliable. The operating model around it has changed.
That is also one of the things worth looking at during a SQL Server Health Check. The existence of an HA configuration tells us relatively little on its own. Its purpose, dependencies, maintenance procedures and current level of operational understanding are much more informative.
Testing a SQL Server high availability design
Documentation and architecture diagrams can describe what should happen during a failure. A controlled failover shows what actually happens.
It verifies whether the expected node or replica can take over, but that is only the beginning of the test. The application also needs to reconnect, monitoring should recognize the role change and required SQL Server Agent jobs and supporting processes need to behave correctly. The team also needs a reliable way to decide whether the service is healthy after the transition.
Testing turns an assumed recovery time into something measurable. It may show that SQL Server changes roles within seconds while the application needs several additional minutes before users can work again. It can expose a DNS dependency, a connection string that still points directly to one server or an operational step that everyone assumed automation would handle.
Those findings are not arguments against the HA technology. They are information about the complete service.
This is why I am cautious with designs that look convincing on paper but that teams rarely exercise. If the team considers failover too risky to test during a controlled maintenance window, relying on the same process during an uncontrolled production incident deserves closer examination.
AG or FCI becomes a useful question after this
Once the failure domain, recovery objectives, maintenance requirements and operating model are clear, comparing SQL Server technologies becomes much more productive.
An FCI can be a natural fit where instance-level continuity is important and the surrounding storage architecture already provides the required availability. An Availability Group can fit a design that needs separate database replicas, database-level failover or replicas across different infrastructure or locations. In some environments, teams deliberately combine the technologies because they address different failure scenarios.
There are further differences around storage, networking, instance-level objects, licensing, readable replicas, administration and SQL Server version capabilities. Those details matter when evaluating a concrete architecture, but they are not a substitute for the requirements that lead to it.
A long AG-versus-FCI feature matrix can otherwise create a false sense of precision. A feature may be technically attractive without solving a problem the workload actually has. Conversely, a comparatively simple design can be completely appropriate when it matches the required failure domain, RPO, RTO and operating capabilities.
Before approving a SQL Server high availability design, I would therefore want to describe several scenarios in ordinary language. We should know what happens when one Windows node disappears, when storage becomes unavailable and when the primary location can no longer be used. For each relevant scenario, we should also understand how much data the business could lose and how long the application may remain unavailable. Finally, the team should have demonstrated the recovery process rather than only documented it.
That gives us something much more useful than starting with the question, “Should we use an Availability Group or an FCI?”
It tells us what the architecture actually needs to accomplish.
Once that is clear, choosing the SQL Server technology becomes a technical decision with defined boundaries rather than the starting point of the design.
Foto von huanyu chen auf Unsplash
