<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>CraftedSQL</title>
	<atom:link href="https://www.craftedsql.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.craftedsql.com/</link>
	<description>Tailored SQL Solutions</description>
	<lastBuildDate>Sat, 26 Sep 2026 14:54:58 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://www.craftedsql.com/wp-content/uploads/2024/11/CraftedSQL-Website-Icon-150x150.png</url>
	<title>CraftedSQL</title>
	<link>https://www.craftedsql.com/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>SQL Server Execution Plans Need Runtime Context</title>
		<link>https://www.craftedsql.com/sql-server-execution-plans-runtime-context/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 10:37:00 +0000</pubDate>
				<category><![CDATA[Performance & Optimization]]></category>
		<category><![CDATA[Cardinality Estimation]]></category>
		<category><![CDATA[Execution Plans]]></category>
		<category><![CDATA[Parameter Sensitivity]]></category>
		<category><![CDATA[Performance Troubleshooting]]></category>
		<category><![CDATA[Query Performance]]></category>
		<category><![CDATA[Query Store]]></category>
		<category><![CDATA[SQL Server]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21189</guid>

					<description><![CDATA[<p>SQL Server execution plans are often the first place I look when a query behaves differently from what I would expect. They tell me how SQL Server chose to access the data, which join strategies it selected and where estimates may not match what actually happened during execution. An actual plan can add runtime information ... <a title="SQL Server Execution Plans Need Runtime Context" class="read-more" href="https://www.craftedsql.com/sql-server-execution-plans-runtime-context/" aria-label="Read more about SQL Server Execution Plans Need Runtime Context">Read more</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-execution-plans-runtime-context/">SQL Server Execution Plans Need Runtime Context</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">SQL Server execution plans are often the first place I look when a query behaves differently from what I would expect. They tell me how SQL Server chose to access the data, which join strategies it selected and where estimates may not match what actually happened during execution. An actual plan can add runtime information as well, so there is usually a lot to work with.</p>



<p class="wp-block-paragraph">The problem starts when people treat the plan as if it were the complete explanation for a performance incident.</p>



<p class="wp-block-paragraph">Suppose users report that an application was slow around 10:00. Someone investigates an hour later, runs the query again and captures an actual execution plan. The query is fast now. The plan looks reasonable. There may still be useful clues in it, but we are already looking at a different execution under different conditions. If I want to understand what happened at 10:00, the plan captured at 11:00 is only one part of the evidence.</p>



<p class="wp-block-paragraph">That distinction sounds obvious, but it matters because execution plans are very persuasive. They are visual, detailed and concrete. It is easy to find an operator that looks suspicious and start tuning from there, even when we have not yet established whether that operator had anything to do with the problem users experienced.</p>



<h2 class="wp-block-heading">What SQL Server execution plans can tell you</h2>



<p class="wp-block-paragraph">An execution plan describes the strategy SQL Server selected for a statement. The optimizer makes that decision using the information available at compile time: statistics, indexes, cardinality estimates, parameter values and other conditions that can influence the plan.</p>



<p class="wp-block-paragraph">An estimated plan shows that compiled strategy without running the query. An actual execution plan adds information from the execution itself, including actual row counts and, depending on the plan and SQL Server version, other runtime details. Microsoft explains the distinction in its documentation on <a href="https://learn.microsoft.com/en-us/sql/relational-databases/performance/display-and-save-execution-plans">displaying and saving execution plans</a>.</p>



<p class="wp-block-paragraph">For troubleshooting, I usually prefer the actual plan because I can compare what SQL Server expected with what it really processed. If SQL Server expected a few rows but processed several hundred thousand instead, that is worth investigating. The mismatch may affect join choices, memory requirements or the amount of work performed further up the plan.</p>



<p class="wp-block-paragraph">But the difference itself is not yet a diagnosis. I still need to know whether it helps explain the slow execution I am investigating.</p>



<h3 class="wp-block-heading">Parameter-sensitive workloads change the picture</h3>



<p class="wp-block-paragraph">This becomes especially relevant with parameter-sensitive workloads. A stored procedure may use the same cached plan for one parameter value that returns ten rows and another that returns several hundred thousand. The plan shape can remain almost unchanged while the actual workload is very different. Data distribution can make that difference larger, and so can changes in the surrounding workload. A plan that works perfectly well for one execution can be much less suitable for another without looking obviously broken when viewed in isolation.</p>



<p class="wp-block-paragraph">The same caution applies in the other direction. A scan, a warning or a high estimated cost may look uncomfortable, but none of those things automatically explains the incident. A scan against a small table may be entirely appropriate. An expensive operation that runs once at night may matter less than a small query executed thousands of times during business hours.</p>



<p class="wp-block-paragraph">What I am trying to establish is not whether the plan contains something that could be improved. On a sufficiently busy SQL Server, I can usually find something. I want to know whether the plan explains the behavior that actually caused the problem.</p>



<h2 class="wp-block-heading">The surrounding workload changes what the plan means</h2>



<p class="wp-block-paragraph">A query can use the same execution plan and still take very different amounts of time.</p>



<p class="wp-block-paragraph">Another transaction may block it. Storage latency may increase. CPU pressure can come from unrelated activity, while several normally harmless queries may start running at the same time and compete for the same resources. None of that necessarily changes the plan shape.</p>



<p class="wp-block-paragraph">This is one reason I do not like looking at SQL Server execution plans completely on their own during a production incident. The plan shows what SQL Server intended to do with the query, but elapsed time also includes waiting. If the query spent most of that time waiting for a lock or for storage, tuning an operator may solve the wrong problem.</p>



<p class="wp-block-paragraph">Wait statistics can help here because they add another view of what SQL Server was waiting for. They have the same limitation, though: they are evidence, not a diagnosis on their own. I wrote about that in <a href="https://www.craftedsql.com/sql-server-wait-statistics-context/">SQL Server Wait Statistics Are Not a Diagnosis</a>.</p>



<p class="wp-block-paragraph">This is also where the timing of the investigation becomes important. If I have an actual plan from the affected execution, together with runtime metrics and information about blocking or resource pressure at the same time, the pieces can start to support each other. If I only have a plan captured later, I have to be much more careful about what I conclude from it.</p>



<p class="wp-block-paragraph">That is often the difference between finding a plausible tuning opportunity and actually explaining an incident.</p>



<h2 class="wp-block-heading">When SQL Server execution plans need history</h2>



<p class="wp-block-paragraph">Teams often investigate SQL Server performance problems after the interesting part has already happened.</p>



<p class="wp-block-paragraph">The application was slow earlier in the morning. Users complained, perhaps monitoring raised an alert, and by the time someone starts looking closely everything appears normal again. CPU is back at its usual level, the blocking chain has disappeared and the query that supposedly caused trouble now finishes quickly.</p>



<p class="wp-block-paragraph">At that point, current SQL Server execution plans can tell me how the query behaves now. What I really need is a way to compare the current behavior with the period where the problem existed.</p>



<h3 class="wp-block-heading">Looking back with Query Store</h3>



<p class="wp-block-paragraph">Query Store helps with exactly that. It keeps query texts, execution plans and aggregated runtime statistics over time, which makes it possible to look back instead of relying only on what I can reproduce at the moment of investigation. Microsoft describes the collection process in <a href="https://learn.microsoft.com/en-us/sql/relational-databases/performance/how-query-store-collects-data?view=sql-server-ver16">How Query Store Collects Data</a>.</p>



<p class="wp-block-paragraph">I may find that the query changed plans shortly before duration increased. That does not prove the new plan caused the regression, but it gives me something concrete to investigate. I can compare both plans, look at the affected runtime intervals and see whether the timing fits.</p>



<p class="wp-block-paragraph">Sometimes the more interesting finding is that the plan did not change at all.</p>



<p class="wp-block-paragraph">If Query Store shows the same plan before, during and after the incident, forcing a different plan becomes much harder to justify. I would start looking more closely at the parameter values used during the slow period, concurrency, blocking or resource pressure. The unchanged plan does not tell me what the cause was, but it removes one attractive explanation from the top of the list.</p>



<p class="wp-block-paragraph">That historical perspective is the reason I consider Query Store so useful in production troubleshooting. I covered that in more detail in <a href="https://www.craftedsql.com/sql-server-query-store-performance-history/">SQL Server Query Store: Why Performance History Matters</a>.</p>



<p class="wp-block-paragraph">Without history, it is very easy to compare a bad user experience from 10:00 with a perfectly healthy execution from 11:00 and assume both represent the same situation.</p>



<p class="wp-block-paragraph">They may not.</p>



<h2 class="wp-block-heading">Before I change production, I want the story to make sense</h2>



<p class="wp-block-paragraph">A performance investigation eventually reaches the point where someone wants to change something. Maybe an index is missing. Maybe a different plan would work better. Perhaps a query needs to be rewritten, or a temporary mitigation is required because the application is already affected.</p>



<p class="wp-block-paragraph">During an active incident, restoring service can take priority. There are situations where I will make a reversible operational change before I understand every detail of the root cause. That is different from deciding what the permanent fix should be.</p>



<p class="wp-block-paragraph">Permanent tuning changes have consequences. An index helps reads but also affects writes and consumes storage. A hint limits the optimizer&#8217;s choices. A forced plan can stabilize a regression and later become a problem itself when the data or workload changes.</p>



<p class="wp-block-paragraph">Before I make that kind of change, I want the evidence to fit together reasonably well. The plan should match the execution I am trying to explain, and the runtime behavior should support the hypothesis. Timing matters too. If historical data is available, I also want to know whether the behavior was unusual or simply normal workload seen at an inconvenient moment.</p>



<p class="wp-block-paragraph">That does not require perfect certainty. Production troubleshooting rarely gives us that.</p>



<p class="wp-block-paragraph">It does require enough context that I can explain why I expect the change to address the actual problem rather than merely improve something that looked suspicious in the plan.</p>



<p class="wp-block-paragraph">For me, that is where SQL Server execution plans are most useful. They narrow the investigation, show me where SQL Server&#8217;s decisions deserve a closer look and give me something concrete to compare against runtime behavior and history.</p>



<p class="wp-block-paragraph">They are strong evidence, but rarely the whole story.</p>



<p class="has-small-font-size wp-block-paragraph">Foto von <a href="https://unsplash.com/de/@davidkristianto?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">David Kristianto</a> auf <a href="https://unsplash.com/de/fotos/luftaufnahme-eines-stark-befahrenen-autobahnkreuzes-in-der-abenddammerung-zzM7voOp6Tk?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-execution-plans-runtime-context/">SQL Server Execution Plans Need Runtime Context</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Azure Arc-enabled SQL Server: Visibility Before Migration</title>
		<link>https://www.craftedsql.com/azure-arc-sql-server-visibility-before-migration/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 13:44:48 +0000</pubDate>
				<category><![CDATA[Hybrid SQL Server]]></category>
		<category><![CDATA[Azure Arc SQL Server]]></category>
		<category><![CDATA[hybrid SQL Server]]></category>
		<category><![CDATA[SQL Server estate management]]></category>
		<category><![CDATA[SQL Server inventory]]></category>
		<category><![CDATA[SQL Server modernization]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21182</guid>

					<description><![CDATA[<p>Azure Arc SQL Server can be useful long before an organization decides to move a database anywhere. In many SQL Server environments, there is an earlier problem to solve: before anyone makes a sensible decision about migration, consolidation or platform strategy, somebody needs to understand what is already there. That sounds obvious, but SQL Server ... <a title="Azure Arc-enabled SQL Server: Visibility Before Migration" class="read-more" href="https://www.craftedsql.com/azure-arc-sql-server-visibility-before-migration/" aria-label="Read more about Azure Arc-enabled SQL Server: Visibility Before Migration">Read more</a></p>
<p>The post <a href="https://www.craftedsql.com/azure-arc-sql-server-visibility-before-migration/">Azure Arc-enabled SQL Server: Visibility Before Migration</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph"><strong>Azure Arc SQL Server</strong> can be useful long before an organization decides to move a database anywhere. In many SQL Server environments, there is an earlier problem to solve: before anyone makes a sensible decision about migration, consolidation or platform strategy, somebody needs to understand what is already there.</p>



<p class="wp-block-paragraph">That sounds obvious, but SQL Server estates rarely grow as neatly as architecture diagrams suggest. A few systems may have been introduced years ago for applications that are still business-critical. Acquisitions, infrastructure changes or temporary projects may have added others along the way. Some run on physical servers, others on virtual machines, and ownership is not always as clear as it once was.</p>



<p class="wp-block-paragraph">The result is not necessarily a badly managed environment. It may simply be an environment that has grown over time and now needs a clearer view.</p>



<h2 class="wp-block-heading">Why visibility should come before migration</h2>



<p class="wp-block-paragraph">At some point, somebody asks a larger question. Which workloads should remain where they are? Which systems need an upgrade? Is there a reason to move anything to Azure? Are there SQL Server versions in use that require attention?</p>



<p class="wp-block-paragraph">Those discussions become difficult when the inventory is incomplete or spread across several sources. As discussed in <a href="/sql-server-workload-placement-on-premises-azure/">Why Workload Placement Matters More Than the Platform Label for SQL Server</a>, a target platform should follow an understanding of the workload and the environment around it, rather than the other way around.</p>



<p class="wp-block-paragraph">This is the part of Azure Arc SQL Server that I find more interesting than the migration story.</p>



<p class="wp-block-paragraph">Azure Arc can bring SQL Server instances running outside Azure into the Azure management plane without moving the databases themselves. SQL Server continues to run on the existing server, whether that is on-premises or elsewhere, while Azure provides a central view of the connected instance. Microsoft describes the current scope in the <a href="https://learn.microsoft.com/en-us/sql/sql-server/azure-arc/overview">Azure Arc-enabled SQL Server overview</a>.</p>



<p class="wp-block-paragraph">That distinction matters because it separates two decisions that often get mixed together: where a workload should run and how the organization wants to manage the SQL Server estate.</p>



<p class="wp-block-paragraph">A SQL Server may have perfectly good reasons to remain on-premises. Latency, application dependencies, licensing, operational constraints or simply the absence of a convincing business case may all support that decision. None of those reasons prevent an organization from improving the way it inventories or governs that server.</p>



<p class="wp-block-paragraph">For a mixed estate, that can already be useful long before migration becomes a serious discussion.</p>



<h2 class="wp-block-heading">Azure Arc SQL Server and the inventory problem</h2>



<p class="wp-block-paragraph">Most organizations can produce a list of SQL Servers somehow. The more interesting question is how much confidence they have in that list.</p>



<p class="wp-block-paragraph">If I am looking at an existing estate, I would want to know more than host names. I would want to understand which SQL Server instances are actually running, which versions and editions they use, which databases belong to them and whether the information is current enough to support a decision.</p>



<p class="wp-block-paragraph">This becomes more important when environments span several locations or operational teams. A CMDB may contain one view, the monitoring platform another, while individual administrators maintain additional information because experience tells them that neither source shows the complete picture.</p>



<p class="wp-block-paragraph">Azure Arc does not magically resolve those organizational differences. What it can provide is another technical source of inventory information based on systems that the organization has actually connected.</p>



<p class="wp-block-paragraph">For SQL Server, this can include information about instances and databases and give teams a common Azure representation for servers that physically remain elsewhere.</p>



<p class="wp-block-paragraph">That is useful, but I would be careful not to confuse inventory with understanding.</p>



<p class="wp-block-paragraph">An old SQL Server version in a portal does not tell you whether the system can simply move to a newer release. A database name does not tell you which application depends on it. And a configuration that differs from a recommendation does not tell you whether changing it would improve the system or create another problem.</p>



<p class="wp-block-paragraph">Those questions still require context. This is also where a broader <a href="/sql-server-health-check-operational-risk/">SQL Server Health Check</a> can become useful: not to create another inventory, but to understand which technical findings have an operational consequence and which ones simply describe the environment as it exists today.</p>



<h2 class="wp-block-heading">Where central visibility becomes useful</h2>



<p class="wp-block-paragraph">The value becomes clearer when inventory information connects to an actual operational problem.</p>



<p class="wp-block-paragraph">Imagine a company with SQL Servers in several locations where lifecycle reviews still depend heavily on manual work. Every few months, somebody collects version information, checks which systems still matter and tries to determine where upgrades deserve attention. A central inventory can remove part of that repeated effort.</p>



<p class="wp-block-paragraph">The same applies when security, infrastructure and database teams need a common view of the environment. They may not need exactly the same level of technical detail, but they do need to be talking about the same systems.</p>



<p class="wp-block-paragraph">Azure Arc can also extend the discussion beyond basic inventory. Microsoft provides additional SQL Server capabilities through the Azure extension for SQL Server, including assessment, management and security-related integrations. The available functions depend on the connected environment, configuration and licensing, so I would treat the current Microsoft documentation as the source of truth rather than assume that every Arc-connected SQL Server offers the same set of capabilities.</p>



<p class="wp-block-paragraph">This is where I would resist the temptation to turn the discussion into a feature checklist. The interesting part is not how many functions the platform offers. It is whether one of them removes a problem that currently costs time, creates uncertainty or leads to weak decisions.</p>



<p class="wp-block-paragraph">If an organization already has reliable inventory, good lifecycle management and established security processes, Arc may add comparatively little in that area. If those processes rely on manual work, differ between teams or leave important gaps, the calculation can look very different.</p>



<h2 class="wp-block-heading">Azure Arc SQL Server also adds an operational layer</h2>



<p class="wp-block-paragraph">There is another part of the architecture that should not disappear behind the promise of central management. Azure Arc is not simply a passive view into an existing SQL Server estate.</p>



<p class="wp-block-paragraph">To use it, an organization connects the servers to Azure. The Azure Connected Machine agent becomes part of the server environment, while SQL Server-specific functions use the Azure extension for SQL Server. Teams need to define permissions, provide the required connectivity and decide who owns the resulting Azure resources and operational processes. Microsoft documents those dependencies in the <a href="https://learn.microsoft.com/en-us/sql/sql-server/azure-arc/prerequisites">prerequisites for Azure Arc-enabled SQL Server</a>.</p>



<p class="wp-block-paragraph">None of this is unusual for an enterprise management platform, but it matters when deciding whether Arc actually makes the environment easier to operate.</p>



<p class="wp-block-paragraph">I would therefore be reluctant to start an Arc project simply because the technology exists or because an organization already uses Azure elsewhere. I would first look at the current operating model and the problems people deal with today.</p>



<p class="wp-block-paragraph">How difficult is it to identify the SQL Servers that exist? How much manual effort goes into lifecycle reviews? Do different teams maintain different inventories? Do important decisions slow down because somebody first has to collect basic information from several places?</p>



<p class="wp-block-paragraph">If those questions expose a real gap, Azure Arc has something concrete to prove itself against. That is much more useful than starting with the product documentation and then looking for a reason to deploy the technology.</p>



<h2 class="wp-block-heading">A hybrid SQL Server estate does not need one destination</h2>



<p class="wp-block-paragraph">There is also a broader architectural point here. Hybrid SQL Server environments are unlikely to disappear simply because cloud platforms continue to mature. Some workloads will move. Teams will replace some systems. Others will remain on SQL Server for a long time because that continues to be the appropriate place for them.</p>



<p class="wp-block-paragraph">Treating every modernization discussion as a migration project therefore creates an unnecessarily narrow choice.</p>



<p class="wp-block-paragraph">A company may decide that several SQL Server workloads should stay exactly where they are while still wanting better inventory, more consistent governance or closer integration with an Azure-based management model. Azure Arc makes that combination possible.</p>



<p class="wp-block-paragraph">It does not decide whether a workload should migrate, and it does not remove the need to understand the applications, dependencies and operational risks behind the SQL Server instances it discovers. What it can do is make the estate easier to see while teams work through those decisions.</p>



<p class="wp-block-paragraph">For me, that is the more useful way to think about Azure Arc-enabled SQL Server. It is not necessarily the first step of a cloud migration. It is one possible way to improve how a distributed SQL Server estate is understood and managed.</p>



<p class="wp-block-paragraph">Sometimes that visibility may eventually support a migration. In other cases, the better outcome is simply having enough information to understand why a workload should stay where it is.</p>



<p class="wp-block-paragraph">If the difficult part is not the technology itself but deciding how an existing SQL Server estate should evolve, the <a href="/sql-server-consulting-services/">SQL Server consulting and modernization work at CraftedSQL</a> focuses on that boundary between infrastructure, workload requirements, operational risk and the next reasonable technical step.</p>



<p class="has-small-font-size wp-block-paragraph">Foto von <a href="https://unsplash.com/de/@imgix?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">imgix</a> auf <a href="https://unsplash.com/de/fotos/serverracks-im-rechenzentrum-klWUhr-wPJ8?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></p>
<p>The post <a href="https://www.craftedsql.com/azure-arc-sql-server-visibility-before-migration/">Azure Arc-enabled SQL Server: Visibility Before Migration</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>A Restored Database Is Not the Same as a Recovered Service</title>
		<link>https://www.craftedsql.com/sql-server-restore-validation/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 10 Sep 2026 11:33:00 +0000</pubDate>
				<category><![CDATA[High Availability & Recovery]]></category>
		<category><![CDATA[Backup and Recovery]]></category>
		<category><![CDATA[Disaster Recovery]]></category>
		<category><![CDATA[Restore Validation]]></category>
		<category><![CDATA[RPO]]></category>
		<category><![CDATA[RTO]]></category>
		<category><![CDATA[SQL Server]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21171</guid>

					<description><![CDATA[<p>A SQL Server restore can complete successfully while the business service behind it remains unavailable. The database may return to an ONLINE state, but that status only tells us something about SQL Server itself. SQL Server restore validation therefore needs to look further: can the application connect, are the required dependencies available, and is the ... <a title="A Restored Database Is Not the Same as a Recovered Service" class="read-more" href="https://www.craftedsql.com/sql-server-restore-validation/" aria-label="Read more about A Restored Database Is Not the Same as a Recovered Service">Read more</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-restore-validation/">A Restored Database Is Not the Same as a Recovered Service</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">A SQL Server restore can complete successfully while the business service behind it remains unavailable. The database may return to an <code>ONLINE</code> state, but that status only tells us something about SQL Server itself. SQL Server restore validation therefore needs to look further: can the application connect, are the required dependencies available, and is the recovered data at the point the business actually needs?</p>



<p class="wp-block-paragraph">If a restore test only needs to prove that SQL Server can restore a backup file, reaching <code>ONLINE</code> may answer the immediate technical question. When the goal is to demonstrate recovery of a business-critical service, however, the database state provides only part of the evidence. Stopping the test at that point can create more confidence than the result really supports.</p>



<p class="wp-block-paragraph">Backup and recovery are often discussed as though restoring the database marks the end of the process. From a SQL Server perspective, that may be true, but operationally the more useful question is what the organization can actually use once the restore has finished.</p>



<h2 class="wp-block-heading">A successful restore proves something important, but not everything</h2>



<p class="wp-block-paragraph">The first task of a restore test is practical: confirm that the available backups can recreate the database. A backup job may run successfully for months and produce valid-looking backup files. Even so, it does not provide the same evidence as restoring those files under controlled conditions.</p>



<p class="wp-block-paragraph">Depending on the recovery requirement, the technical validation may also confirm the expected backup chain and check the resulting database state. It may also need to establish which point in time the restore has recovered. Microsoft documents the <a href="https://learn.microsoft.com/en-us/sql/relational-databases/backup-restore/restore-and-recovery-overview-sql-server" rel="noopener">SQL Server restore and recovery process</a> in detail, although those technical steps still answer questions about the database rather than the service that depends on it.</p>



<p class="wp-block-paragraph">The operational part begins once SQL Server has completed its work, because an application rarely depends on a restored database alone. It also needs to find the SQL Server instance, authenticate correctly and interact with the other components the workload requires.</p>



<p class="wp-block-paragraph"><strong>That is where the difference between database restored and service recovered becomes visible.</strong></p>



<h2 class="wp-block-heading">The database is part of a larger dependency chain</h2>



<p class="wp-block-paragraph">Some dependencies an application needs do not live inside the restored database at all. SQL Server logins, SQL Agent jobs, credentials, linked servers and instance-level configuration may matter. Application configuration, certificates, DNS, network access or external services can introduce further dependencies outside SQL Server.</p>



<p class="wp-block-paragraph">This does not mean that every restore test should become an exhaustive inventory of every possible server object or infrastructure component. The useful question is narrower: which dependencies does this particular service require, and does the recovery design include them?</p>



<p class="wp-block-paragraph">During normal operation, many of those relationships remain largely invisible because everything is already in place. A recovery makes them much easier to see, especially when the database is technically available but the application still cannot use it.</p>



<p class="wp-block-paragraph">Restore testing therefore provides value beyond validating the backup itself. It also shows how well the organization understands the system and the dependencies needed to bring it back into use.</p>



<h2 class="wp-block-heading">Application connectivity deserves its own validation</h2>



<p class="wp-block-paragraph">A common gap appears when the database team checks the restored database, but nobody tests the application against the recovered environment. From the SQL Server side, everything may look reasonable: the database is online, queries run and the expected objects exist. None of those checks proves that the application can establish its own connection and perform the work for which the database exists.</p>



<p class="wp-block-paragraph">Authentication may behave differently on the recovery system. A connection string may still reference another server, a required login may be missing or mapped incorrectly, or a network dependency may differ from production. In each case, the restore can succeed technically while the service remains unusable.</p>



<p class="wp-block-paragraph">For a simple application, proving service availability may require little more than connecting through the application and executing a known function. In a more complex environment, the database team may not be able to make that judgement alone. Application support or a business owner may need to confirm that the relevant workflow behaves as expected.</p>



<p class="wp-block-paragraph">The point is not to prescribe one universal test. Someone needs to validate the recovered environment from the perspective of the service, rather than only from the perspective of SQL Server.</p>



<h2 class="wp-block-heading">The right database can still contain the wrong data</h2>



<p class="wp-block-paragraph">Another distinction appears when the database itself is healthy, but the restored point in time does not meet the business requirement.</p>



<p class="wp-block-paragraph">A database can restore successfully and pass technical checks while still leaving the organization with more data loss than it intended to accept. Suppose the most recent usable recovery point is several hours old, while the agreed recovery point objective allows only a much smaller gap. In that case, the restore worked, but the recovery objective did not.</p>



<p class="wp-block-paragraph">This is where RPO becomes operational rather than theoretical. The relevant question is not simply whether the organization can restore data. It also needs to know which data would actually be available after a realistic failure and whether that recovery point is acceptable for the affected process.</p>



<p class="wp-block-paragraph">Backup frequency, transaction log backups, retention and the availability of the required backup chain all influence that result. Effective SQL Server restore validation should therefore make the recovered point in time visible. The database state alone does not prove that the recovery requirement has been met.</p>



<p class="wp-block-paragraph">Otherwise, the technical team may correctly report that the restore succeeded, while the application owner later discovers that the available data is older than expected.</p>



<h2 class="wp-block-heading">Recovery time does not stop when RESTORE finishes</h2>



<p class="wp-block-paragraph">The same distinction applies to recovery time, because the SQL Server restore represents only one part of the time needed to return a service to use. A useful SQL Server restore validation process should therefore measure more than the restore operation itself. Application checks, dependency validation and operational approval may add significant time before the service can resume.</p>



<p class="wp-block-paragraph">A restore operation is easy to measure because it has a clear start and finish. Suppose the database takes 35 minutes to restore, but the teams then need another hour to resolve dependencies and reconnect the application. They may also need to validate the data and approve the return to service. In that case, 35 minutes remains useful technical information, but it is not the effective service recovery time.</p>



<p class="wp-block-paragraph">For business continuity planning, the meaningful endpoint is the moment when the required service can actually resume. Reaching that point may involve several teams. Infrastructure may provide the platform, the DBA may restore and validate the database, and application support may verify connectivity and behaviour. The business owner may still need to decide whether the recovered state is acceptable.</p>



<p class="wp-block-paragraph">When teams have never tested those handoffs, the organization may know how long SQL Server needs for a restore. It still may not know how long the complete recovery really takes.</p>



<h2 class="wp-block-heading">Recovery needs a defined end condition</h2>



<p class="wp-block-paragraph">That leads to one of the less technical questions in the process: who decides that recovery is complete?</p>



<p class="wp-block-paragraph">The DBA can confirm a successful database restore and verify that SQL Server sees the database as available. Infrastructure can confirm that the server and network work as expected, while application support can verify connectivity and application behaviour. Depending on the system, a business owner may also need to confirm that the recovered data and the relevant workflow are acceptable.</p>



<p class="wp-block-paragraph">No universal division of responsibility fits every environment, nor does one need to. The boundary should simply be understood before an incident occurs. If every team considers its own technical task complete without a shared definition of the final outcome, several successful individual steps can still leave the service unavailable beyond the required recovery time.</p>



<p class="wp-block-paragraph">A useful restore test therefore needs an operational end condition. Instead of ending with “the database is ONLINE,” the test should establish that the required data is available and the application can use it. The relevant dependencies also need to be in place, and the responsible people need enough evidence to return the service to operation.</p>



<p class="wp-block-paragraph">That provides a different level of assurance, because it describes an operational outcome rather than a database state.</p>



<h2 class="wp-block-heading">What SQL Server restore validation should actually prove</h2>



<p class="wp-block-paragraph">Backup strategy and restore validation are sometimes treated as separate administrative topics, although they address the same recovery question. Backup frequency, retention, protected storage, monitoring and transaction log handling exist because the organization expects to recover something from them when needed.</p>



<p class="wp-block-paragraph">Restore testing allows the organization to test those assumptions against an actual outcome.</p>



<p class="wp-block-paragraph">Can the team locate and access the required backups? Does the necessary backup chain exist, and can the team restore the database away from the original system if necessary? Are the important dependencies known? Can the application use the recovered database, and does the resulting recovery point meet the requirement?</p>



<p class="wp-block-paragraph">Those questions provide more useful evidence than a green backup job alone. They connect the technical backup process to the service it is meant to protect.</p>



<p class="wp-block-paragraph">The earlier CraftedSQL article <a href="/sql-server-backup-strategy/" target="_blank" rel="noreferrer noopener">SQL Server Backup Strategy: Why Successful Backups Are Not Enough</a> looked at this from the backup side. Restore validation provides the corresponding recovery perspective: the backup process becomes much more meaningful once the organization has evidence of what it can actually recover from it.</p>



<h2 class="wp-block-heading">The test should reflect the service you are trying to recover</h2>



<p class="wp-block-paragraph">No single restore-validation procedure fits every SQL Server environment, because recovery requirements differ substantially. A reporting database that may remain unavailable for several hours does not need the same validation process as an operational system with stricter recovery requirements. A self-contained application also has a different dependency chain from a service that spans several databases, interfaces and infrastructure components.</p>



<p class="wp-block-paragraph">Good SQL Server restore validation therefore starts with the expected recovery outcome rather than with a generic checklist. What must be available again, and how much data loss can the organization accept? Which dependencies are required? Who can verify that the application works, and who is responsible for releasing the service back into operation?</p>



<p class="wp-block-paragraph">Once those questions are clear, the team can test the restore process against them. That makes it much easier to define what a successful recovery test actually means.</p>



<p class="wp-block-paragraph">A database reaching <code>ONLINE</code> remains an important milestone, but it is not the same thing as having recovered the service that depends on it.</p>



<p class="has-small-font-size wp-block-paragraph">Foto von <a href="https://unsplash.com/de/@hyoshining?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Hyoshin Choi</a> auf <a href="https://unsplash.com/de/fotos/eine-person-die-auf-einer-brucke-geht-bN4kaXPNxPA?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></p>



<p class="wp-block-paragraph"></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-restore-validation/">A Restored Database Is Not the Same as a Recovered Service</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>When Permissions Outlive Their Context in SQL Server</title>
		<link>https://www.craftedsql.com/sql-server-permissions-historical-access/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 03 Sep 2026 16:27:05 +0000</pubDate>
				<category><![CDATA[Consulting & Practice]]></category>
		<category><![CDATA[Service Accounts]]></category>
		<category><![CDATA[SQL Server]]></category>
		<category><![CDATA[SQL Server Permissions]]></category>
		<category><![CDATA[SQL Server Security]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21143</guid>

					<description><![CDATA[<p>Most SQL Server permissions begin with a perfectly reasonable decision. An application needs access to another database, a service account requires additional rights for a deployment, or an administrator receives broader permissions while a migration is in progress. At that point, the reason is probably known, and there may even be a ticket, a project ... <a title="When Permissions Outlive Their Context in SQL Server" class="read-more" href="https://www.craftedsql.com/sql-server-permissions-historical-access/" aria-label="Read more about When Permissions Outlive Their Context in SQL Server">Read more</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-permissions-historical-access/">When Permissions Outlive Their Context in SQL Server</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Most SQL Server permissions begin with a perfectly reasonable decision. An application needs access to another database, a service account requires additional rights for a deployment, or an administrator receives broader permissions while a migration is in progress.</p>



<p class="wp-block-paragraph">At that point, the reason is probably known, and there may even be a ticket, a project document or someone who can explain exactly why the access is required. The problem tends to appear much later, after applications have changed, projects have ended and the people involved have moved to different roles, while the permissions created along the way have remained untouched.</p>



<p class="wp-block-paragraph">The SQL Server still works, the application continues to run and nobody reports a security problem, so there is little operational pressure to revisit those decisions. Over time, however, the environment accumulates access that may still be technically valid while the context behind it slowly disappears.</p>



<p class="wp-block-paragraph">That is where permission reviews become more interesting than simply finding accounts with elevated privileges.</p>



<h2 class="wp-block-heading">The permission itself does not explain why it exists</h2>



<p class="wp-block-paragraph">The SQL Server permissions model gives us enough metadata to build a reasonably detailed picture of access across an instance. We can inspect server and database principals, role memberships, direct grants and Windows groups, then identify accounts that appear unusually powerful or whose purpose is unclear. Microsoft documents the relevant principals, roles and catalog views in its <a href="https://learn.microsoft.com/en-us/sql/relational-databases/security/authentication-access/getting-started-with-database-engine-permissions?view=sql-server-ver17">Database Engine permissions documentation</a>.</p>



<p class="wp-block-paragraph">That inventory is useful, but it is only the starting point because the technical configuration rarely tells the whole story. A login may have broad access because a current application genuinely needs it, or because somebody granted those permissions during a migration six years ago and nobody ever revisited the decision afterwards.</p>



<p class="wp-block-paragraph">From the permission alone, those situations can look almost identical.</p>



<p class="wp-block-paragraph">When I review an established SQL Server environment, I therefore want to understand more than which login belongs to which role. I want to know what uses the access today, who owns the corresponding application or process, and whether the current privilege level still matches the technical requirement.</p>



<p class="wp-block-paragraph">In environments that have been running for many years, answering those questions can be considerably harder than producing the permission inventory itself. SQL Server permissions have a habit of surviving organizational and technical changes much longer than the documentation that once explained them.</p>



<h2 class="wp-block-heading">Reviewing old SQL Server permissions needs context, not assumptions</h2>



<p class="wp-block-paragraph">The age of a login or permission is useful information, but it is not evidence that the access is obsolete.</p>



<p class="wp-block-paragraph">An account may belong to an application that was retired years ago, in which case removing it may be straightforward once that dependency has been confirmed. Another account may appear equally old and equally unfamiliar, yet still belong to an integration that runs once a month, a maintenance process that is used only occasionally, or a component that becomes relevant during recovery.</p>



<p class="wp-block-paragraph">A short observation period will not necessarily reveal those dependencies, and neither will the fact that nobody immediately recognizes the account name.</p>



<p class="wp-block-paragraph">That creates an uncomfortable but important distinction. “Nobody knows what this account does” should trigger an investigation, but it should not automatically trigger a <code>DROP LOGIN</code>, just as uncertainty should not become a reason to preserve unexplained privileged access indefinitely.</p>



<p class="wp-block-paragraph">Both approaches create risk.</p>



<p class="wp-block-paragraph">Leaving unnecessary administrative access in place increases the security exposure of the environment, while removing access without understanding the dependency can turn a security cleanup into an availability incident. A useful permission review therefore has to deal with both sides of that problem rather than optimizing for the shortest possible findings list.</p>



<h2 class="wp-block-heading">Reconstructing the dependency before changing access</h2>



<p class="wp-block-paragraph">Once an unusual permission has been identified, I would start reconstructing the context around it. Is the account associated with a person, an application, a SQL Server Agent job, a Windows service or another technical process? Is there evidence that the access is still used, and can somebody identify an owner who is able to confirm what that process actually requires?</p>



<p class="wp-block-paragraph">If the original documentation is missing, the answers may have to come from several places, including job ownership, service configurations, Windows group memberships, application settings or authentication activity. The exact investigation depends on the environment, but the objective remains the same: gather enough evidence to distinguish required access from access that merely survived.</p>



<p class="wp-block-paragraph">An account that cannot yet be classified should remain visible as an unresolved finding. Lack of context is itself an operational problem because it means that nobody currently has enough information to make a safe decision about that part of the system.</p>



<p class="wp-block-paragraph">In practice, permissions are often only one part of a broader <a href="https://www.craftedsql.com/sql-server-consulting-services/">SQL Server security review</a>. Service accounts and administrative access can expose the same underlying problem from another direction: access still exists, but the reason for it is no longer sufficiently understood.</p>



<h2 class="wp-block-heading">Why service accounts complicate SQL Server permissions reviews</h2>



<p class="wp-block-paragraph">Service accounts are a good example because they can remain unchanged for much longer than the people who originally configured them.</p>



<p class="wp-block-paragraph">An account may have been created during the first deployment of an application, received additional privileges during a later migration and then continued running quietly through several upgrades. Years later, the application is still business-critical, but nobody can explain whether all of the SQL Server permissions accumulated by its service account are still necessary.</p>



<p class="wp-block-paragraph">Sometimes the answer is straightforward once the application owner becomes involved. In other cases, the account has become part of a chain of technical dependencies that crosses SQL Server, Windows services, scheduled tasks and application components, which makes a seemingly simple permission change much harder to evaluate safely.</p>



<p class="wp-block-paragraph">This is why the privilege level alone does not tell me enough.</p>



<p class="wp-block-paragraph">If a service account has broad access, I want to know what would stop working if that access were reduced and whether the answer is based on current knowledge rather than on fear of touching an old system. Microsoft also recommends running SQL Server services with the lowest possible user rights and avoiding unnecessary additional permissions for service accounts; the details depend on the service and account type, but the underlying principle is clear. See <a href="https://learn.microsoft.com/en-us/sql/database-engine/configure-windows/configure-windows-service-accounts-and-permissions?view=sql-server-ver17">Configure Windows Service Accounts and Permissions</a> for the current Microsoft guidance.</p>



<p class="wp-block-paragraph">If nobody can explain what the service account actually requires, the permission problem is accompanied by a second problem: the organization has lost part of the operational knowledge needed to change the system safely. That does not make the current access acceptable, but it changes the next step because the dependency first has to be understood well enough that the change can be controlled.</p>



<h2 class="wp-block-heading">Least privilege is also a production change</h2>



<p class="wp-block-paragraph">The principle of least privilege is simple enough: an account should have the permissions it requires and no more. Applying that principle to a production environment that has grown organically for ten or fifteen years is less simple because reducing permissions changes something that another component may depend on.</p>



<p class="wp-block-paragraph">For significant permissions, I therefore treat remediation like any other production change. Before removing or reducing access, I want to understand the expected dependency, know how the result will be validated and have a realistic rollback path if the assumption turns out to be wrong.</p>



<p class="wp-block-paragraph">The amount of effort should match the possible impact. Removing an obsolete user from an unused development database does not require the same level of preparation as changing the account behind a business-critical application, but both changes should be based on more than the appearance of a permission in a report.</p>



<p class="wp-block-paragraph">This is also where some security reviews become less useful than they could be.</p>



<p class="wp-block-paragraph">A large findings list creates visible activity, yet it does not necessarily tell the organization what should happen next. If twenty accounts are marked because they are old, another ten because they have direct grants and several service accounts because they appear highly privileged, the technical observations still need to be translated into priorities.</p>



<p class="wp-block-paragraph">An undocumented account with server-wide administrative privileges deserves different attention from an old database user with limited access to a non-production database. Likewise, a service account supporting a critical workload may require more careful investigation than several obviously obsolete logins whose dependencies have already been ruled out.</p>



<p class="wp-block-paragraph">Age matters, privilege matters and scope matters, but so do workload importance, current usage and the ability to explain why the access still exists.</p>



<h2 class="wp-block-heading">A useful review should lead to decisions</h2>



<p class="wp-block-paragraph">At the end of a SQL Server permissions review, I do not expect every account to fall into the same category.</p>



<p class="wp-block-paragraph">Some permissions will be appropriate and can remain unchanged. Others will still be required, although the current privilege level is broader than necessary and should be reduced. Some accounts will turn out to be obsolete and can be prepared for removal, while another group will require further investigation because the available evidence is not yet good enough for a safe decision.</p>



<p class="wp-block-paragraph">That last group should not be hidden simply because it makes the review look unfinished. In a mature environment, identifying where operational knowledge is missing can be just as useful as identifying an excessive grant, because it tells us where future changes are likely to carry unnecessary uncertainty.</p>



<p class="wp-block-paragraph">For the important accounts, the review should eventually make three things clear: why the access exists, whether the current permissions are still justified, and who is responsible for that decision.</p>



<p class="wp-block-paragraph">Once those questions can be answered, changing the technical configuration is usually the easier part.</p>



<p class="wp-block-paragraph">This is why historically grown permissions concern me more than many newly granted ones. A recent change normally still has context around it, whether that is a ticket, a project, a deployment or simply somebody who remembers why the decision was made.</p>



<p class="wp-block-paragraph">Older permissions may have survived several generations of applications, administrators and operating procedures while continuing to work exactly as they always did.</p>



<p class="wp-block-paragraph">The access did not disappear.</p>



<p class="wp-block-paragraph">The knowledge around it did.</p>



<p class="wp-block-paragraph">And when privileged access can no longer be explained, that is a good reason to investigate it before either blindly accepting it or blindly removing it.</p>



<p class="has-small-font-size wp-block-paragraph">Foto von <a href="https://unsplash.com/de/@toolmash?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Toolmash Expo</a> auf <a href="https://unsplash.com/de/fotos/elektriker-pruft-schalttafel-mit-multimeter-PkHf7BUWbtk?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-permissions-historical-access/">When Permissions Outlive Their Context in SQL Server</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Why Workload Placement Matters More Than the Platform Label for SQL Server</title>
		<link>https://www.craftedsql.com/sql-server-workload-placement-on-premises-azure/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 27 Aug 2026 09:55:00 +0000</pubDate>
				<category><![CDATA[Architecture & Setup]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21105</guid>

					<description><![CDATA[<p>When companies start discussing SQL Server workload placement, the conversation often begins with the platform. Should the workload move to Azure? Would Azure SQL Managed Instance be a better fit? Should the database stay on-premises for another few years, or would that simply postpone an inevitable migration? Those are reasonable questions, but in my experience ... <a title="Why Workload Placement Matters More Than the Platform Label for SQL Server" class="read-more" href="https://www.craftedsql.com/sql-server-workload-placement-on-premises-azure/" aria-label="Read more about Why Workload Placement Matters More Than the Platform Label for SQL Server">Read more</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-workload-placement-on-premises-azure/">Why Workload Placement Matters More Than the Platform Label for SQL Server</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">When companies start discussing <strong>SQL Server workload placement</strong>, the conversation often begins with the platform. Should the workload move to Azure? Would Azure SQL Managed Instance be a better fit? Should the database stay on-premises for another few years, or would that simply postpone an inevitable migration?</p>



<p class="wp-block-paragraph">Those are reasonable questions, but in my experience they often come too early.</p>



<p class="wp-block-paragraph">Before I can say anything useful about the target platform, I first need to understand the workload itself. What does the application depend on? How is it operated today? Where does the data need to be located? Which availability and recovery requirements exist? And, perhaps most importantly, what problem is the company actually trying to solve by changing the platform?</p>



<p class="wp-block-paragraph">Moving a SQL Server workload to the cloud is not modernization by itself. At the same time, keeping it on-premises does not automatically mean that an environment is outdated. The useful discussion starts somewhere else: with the workload, its dependencies and the operating model around it.</p>



<h2 class="wp-block-heading">SQL Server workload placement starts with the reason for change</h2>



<p class="wp-block-paragraph">A migration project is much easier to evaluate when there is a clear reason behind it. Maybe the existing hardware is approaching end of life, the infrastructure team wants to reduce the amount of hardware it operates itself, or the current platform no longer provides the flexibility or availability the business needs. Those are concrete problems a team can measure and discuss.</p>



<p class="wp-block-paragraph">The situation becomes much harder when the starting point is simply that an existing SQL Server runs on-premises and therefore appears to require modernization. If the current operating model is stable, understood and economically reasonable, moving the workload somewhere else does not automatically make it better.</p>



<p class="wp-block-paragraph">This is also why discussions about cloud adoption sometimes become more complicated than expected. Microsoft recently asked partners what keeps customers from migrating SQL Server workloads to Azure. There are many possible answers to that question, but from my perspective there is an even more useful one to ask first: would migration solve a problem that actually matters to this customer?</p>



<p class="wp-block-paragraph">Without that context, technical feasibility only tells part of the story. Microsoft maintains a useful overview of the different <a href="https://learn.microsoft.com/en-us/azure/azure-sql/azure-sql-iaas-vs-paas-what-is-overview?view=azuresql" target="_blank" rel="noopener">Azure SQL deployment options</a>, but choosing between them only becomes meaningful once the workload requirements are clear.</p>



<h2 class="wp-block-heading">A SQL Server database rarely exists on its own</h2>



<p class="wp-block-paragraph">Looking only at the database engine can make a migration appear simpler than it really is. The instance may use supported features, its size may be manageable, and an assessment may show no major blockers for Azure SQL Managed Instance, Azure SQL Database or a SQL Server virtual machine.</p>



<p class="wp-block-paragraph">That is useful information, but the database is usually only one part of the workload.</p>



<p class="wp-block-paragraph">Applications may depend on other systems in the same network. Jobs may access local file shares. Reporting systems may read from several nearby data sources. Third-party software may require a particular SQL Server version, authentication model or operating system configuration. Monitoring, backup, administrative tooling and support processes may have grown around the current environment over many years.</p>



<p class="wp-block-paragraph">None of these dependencies automatically prevents a migration. They simply change the question. Instead of asking only whether the database can be moved, I want to understand what happens to everything around it if we do. Sometimes that leads to the conclusion that the database and the surrounding application stack should move together. In other cases it exposes dependencies that make a staged or hybrid approach more sensible.</p>



<p class="wp-block-paragraph">This is also why I treat migration planning as an architectural and operational exercise rather than only a database move. The same principle is part of my <a href="https://www.craftedsql.com/sql-server-consulting-services/">SQL Server migration and platform consulting</a>: the target platform should fit the workload, its dependencies and the way the environment will operate afterwards.</p>



<h3 class="wp-block-heading">Latency and data movement can change the decision</h3>



<p class="wp-block-paragraph">Latency and data movement are good examples. An application and its SQL Server database may work perfectly well when both are located in the same data center. If only the database moves to Azure while the application and several dependent systems remain on-premises, that relationship changes.</p>



<p class="wp-block-paragraph">For some workloads this barely matters. Others perform a large number of small synchronous database calls and react badly to additional network latency. Large or frequent data transfers can create another dependency on the network architecture, especially if systems continue to exchange substantial amounts of data between on-premises infrastructure and Azure.</p>



<p class="wp-block-paragraph">The point is not that these workloads should not move. The point is that the database cannot be evaluated in isolation.</p>



<h2 class="wp-block-heading">Availability and recovery still need to be designed</h2>



<p class="wp-block-paragraph">Cloud platforms can make some availability scenarios considerably easier to implement and operate. Managed database services can also remove infrastructure responsibilities that would otherwise remain with an internal team.</p>



<p class="wp-block-paragraph">That can be a very real benefit, but it does not remove the need to understand the actual requirement. How much downtime can the business tolerate? How much data can it afford to lose? What happens when an entire region becomes unavailable? How does the application reconnect after a failure, and who is responsible for testing the recovery process?</p>



<p class="wp-block-paragraph">These questions exist regardless of whether SQL Server runs in a local data center, in an Azure virtual machine or on a managed database platform. The available technical options change, but the business requirement does not disappear.</p>



<p class="wp-block-paragraph">The same is true for backup and recovery. Automated backups and managed infrastructure reduce operational work, but the organization still needs to know what it has to restore, how quickly the service needs to be available again and whether the selected platform can actually meet that requirement.</p>



<p class="wp-block-paragraph">A platform can simplify the implementation. It cannot define the requirement for you.</p>



<h2 class="wp-block-heading">The operating model matters in SQL Server workload placement</h2>



<p class="wp-block-paragraph">One of the strongest arguments for moving a workload to a managed service is not that the database will run in Azure. It is that the responsibility model changes.</p>



<p class="wp-block-paragraph">With a traditional SQL Server installation, someone has to operate the operating system, storage, patching, backups, monitoring, SQL Server configuration and the infrastructure underneath it. Depending on the architecture, high availability and disaster recovery add another layer of work.</p>



<p class="wp-block-paragraph">A managed service can remove part of that responsibility. Azure SQL Database moves much more of the infrastructure and platform operation to Microsoft. Azure SQL Managed Instance keeps more SQL Server compatibility while still reducing a substantial amount of platform administration. A SQL Server running in an Azure virtual machine, on the other hand, leaves many of the familiar responsibilities with the customer.</p>



<p class="wp-block-paragraph">That distinction is far more useful than simply comparing “cloud” with “on-premises”. If an organization is struggling because it no longer wants to maintain database infrastructure, then reducing operational responsibility can be a very good reason to move. If the same organization takes an existing SQL Server virtual machine and recreates almost the same architecture in Azure, the operating model may change much less than expected.</p>



<p class="wp-block-paragraph">That does not make the migration wrong. It just means the benefit has to come from somewhere else.</p>



<h3 class="wp-block-heading">Existing skills and operational capacity matter</h3>



<p class="wp-block-paragraph">The existing team also matters. A company with experienced infrastructure and SQL Server staff may already have mature processes for patching, monitoring, backup, recovery and high availability. In that environment, keeping certain workloads on-premises can remain predictable and economically reasonable.</p>



<p class="wp-block-paragraph">Another company may be in the opposite situation. Hardware maintenance, Windows patching and SQL Server administration may depend on a very small team that is already overloaded. There, reducing the amount of infrastructure the organization has to operate can be far more valuable than preserving an existing architecture.</p>



<p class="wp-block-paragraph">The same workload can therefore lead to different platform decisions in different organizations. This is also why monitoring alone is not enough to judge whether an environment needs a platform change. I discussed that distinction in more detail in <a href="https://www.craftedsql.com/sql-server-monitoring-is-not-diagnosis/">SQL Server Monitoring Is Not Diagnosis</a>.</p>



<h2 class="wp-block-heading">Governance and cost need the same level of detail</h2>



<p class="wp-block-paragraph">Data sovereignty, compliance and governance are frequently mentioned in cloud discussions, but they are only useful decision criteria when they are made specific. Which data is affected? Which regulation, contract or internal policy applies? Where may the organization store the data? Who may administer the platform? Which audit evidence does the organization need?</p>



<p class="wp-block-paragraph">Once the organization answers those questions, it can evaluate the available options properly. For many workloads, Azure can meet demanding compliance and governance requirements. In other cases, contractual obligations, internal policies or dependencies on other systems still influence where the workload can reasonably run.</p>



<p class="wp-block-paragraph">“Compliance” should therefore neither be used as a vague argument against cloud adoption nor dismissed as simple resistance to change. It is a requirement that needs the same level of clarification as any other.</p>



<h3 class="wp-block-heading">Cost needs to reflect the actual operating model</h3>



<p class="wp-block-paragraph">Cost deserves the same treatment. A simple comparison between server hardware and a monthly Azure estimate is rarely sufficient.</p>



<p class="wp-block-paragraph">An on-premises SQL Server has hardware, storage, licensing, backup, power, cooling, lifecycle and operational costs. Azure has compute, storage, backup, networking and service costs, and the licensing model depends heavily on the chosen architecture and on existing entitlements.</p>



<p class="wp-block-paragraph">The difficulty is that some costs are highly visible while others already sit inside existing infrastructure and teams. A company with available virtualization capacity, experienced staff and established processes will calculate very differently from an organization facing a large hardware refresh and lacking the people to operate the platform properly.</p>



<p class="wp-block-paragraph">That is why generic statements about one model always being cheaper than the other are rarely useful. A meaningful comparison has to describe the actual environment before the change and the operating model that will exist afterwards.</p>



<h2 class="wp-block-heading">Hybrid can be a deliberate workload placement decision</h2>



<p class="wp-block-paragraph">Another assumption I would question is that a hybrid architecture is always an intermediate state. Sometimes it is. A company may move workloads in stages and gradually reduce its local infrastructure.</p>



<p class="wp-block-paragraph">But there are also environments where different workloads simply have different requirements. A relatively independent application might be an excellent candidate for Azure SQL Database. Another workload may fit Azure SQL Managed Instance because it still depends on SQL Server features that make Azure SQL Database impractical. A third system may remain on-premises because it has strong local dependencies or because moving it currently produces no meaningful operational benefit.</p>



<p class="wp-block-paragraph">That is not necessarily an unfinished migration. It can simply be the result of treating workloads individually instead of forcing every database into the same target architecture.</p>



<p class="wp-block-paragraph">Microsoft&#8217;s <a href="https://learn.microsoft.com/en-us/sql/sql-server/migrate/?view=sql-server-ver17" target="_blank" rel="noopener">SQL Server migration documentation</a> reflects that broader range of options as well, covering both on-premises SQL Server targets and Azure SQL destinations rather than treating migration as a single predefined path.</p>



<p class="wp-block-paragraph">For me, that is usually the more useful way to think about modernization.</p>



<p class="wp-block-paragraph">I do not consider the physical location of SQL Server to be a meaningful measure of how modern an environment is. I have seen well-operated on-premises systems with clear responsibilities, tested recovery procedures, appropriate monitoring and sensible lifecycle management. And an environment does not become well designed simply because the same problems now run on infrastructure hosted somewhere else.</p>



<p class="wp-block-paragraph">A modernization project should improve something that matters. It may reduce operational responsibility, improve recovery capabilities, remove aging infrastructure, make scaling easier or allow the application architecture to evolve. Azure SQL can be a very good answer to those requirements.</p>



<p class="wp-block-paragraph">In other situations, modernizing the existing SQL Server platform on-premises may be the more sensible next step. Quite often, the answer will be a combination of both.</p>



<p class="wp-block-paragraph">The platform decision becomes much easier once the order of the questions is right.</p>



<p class="wp-block-paragraph"><strong>Understand the workload, its dependencies and the operating model first. Then decide where it should run.</strong></p>



<p class="has-small-font-size wp-block-paragraph">Foto von <a href="https://unsplash.com/de/@crankeli?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Alexander ‎</a> auf <a href="https://unsplash.com/de/fotos/eine-bahnstrecke-die-durch-ein-landliches-gebiet-verlauft-8rKGiR0t5Z0?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-workload-placement-on-premises-azure/">Why Workload Placement Matters More Than the Platform Label for SQL Server</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>SQL Server Health Checks: Small Operational Gaps That Become Expensive Later</title>
		<link>https://www.craftedsql.com/sql-server-health-check-operational-risk/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 20 Aug 2026 08:55:00 +0000</pubDate>
				<category><![CDATA[DBA Essentials]]></category>
		<category><![CDATA[Operational Risk]]></category>
		<category><![CDATA[SQL Server Health Check]]></category>
		<category><![CDATA[SQL Server Operations]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21090</guid>

					<description><![CDATA[<p>A SQL Server environment can run reliably for years while operational gaps remain unnoticed. A structured Health Check helps identify which findings matter, understand their consequences and set meaningful priorities.</p>
<p>The post <a href="https://www.craftedsql.com/sql-server-health-check-operational-risk/">SQL Server Health Checks: Small Operational Gaps That Become Expensive Later</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">A SQL Server Health Check can be useful even when the environment is not failing or showing an obvious problem. Many systems run reliably for years. Applications are available, scheduled jobs complete, backup processes run as expected and there may be no noticeable performance issue. From day to day, there is little reason to question a system that appears to be doing exactly what it should, especially when nobody is currently reporting a problem.</p>



<p class="wp-block-paragraph">That is also why some of the most useful findings during a SQL Server Health Check are not failures at all. They are often things that still work, but whose original context has gradually disappeared. SQL Server environments change over time: teams add databases, replace applications, adjust maintenance routines and move infrastructure to new platforms. Administrators leave, responsibilities change and temporary solutions occasionally remain in place much longer than anyone originally expected.</p>



<p class="wp-block-paragraph">A configuration setting that once solved a specific problem may still exist years later. A SQL Agent job may continue to run every night even though the application it was created for has changed substantially. Somewhere else, a dependency may never have made it into the documentation because everyone involved knew about it at the time. None of this means the environment is badly operated, but a system can still be stable while parts of its operation are no longer fully understood.</p>



<h2 class="wp-block-heading">What a SQL Server Health Check can reveal about stability</h2>



<p class="wp-block-paragraph">When the surrounding conditions remain similar, a SQL Server environment can continue to behave predictably for a long time. The same applications connect to it, jobs run at familiar times and the infrastructure underneath the workload changes only gradually. Small inconsistencies can remain invisible because nothing forces the team to look at them.</p>



<p class="wp-block-paragraph">The situation becomes more interesting when something has to change. An upgrade may require the team to understand why someone set a particular database option years ago. During a migration, an application may suddenly reveal that it depends on a local path, a server name or another component that nobody considered part of the migration. This is also why <a href="https://www.craftedsql.com/sql-server-consulting-services/">SQL Server upgrade and migration planning</a> should start with understanding the existing environment rather than only the target platform. High availability provides another good example. An Availability Group may look perfectly healthy until the first real failover exposes an application dependency outside the database layer.</p>



<p class="wp-block-paragraph">At that point, the technical problem is only part of the work. Someone also has to reconstruct how the environment was supposed to function, often while there is already time pressure. This is where knowing the history of a system becomes operationally useful.</p>



<h2 class="wp-block-heading">Unusual does not automatically mean wrong</h2>



<p class="wp-block-paragraph">One thing I would avoid during a SQL Server Health Check is treating every deviation from a recommendation as something that needs to be fixed. Real environments rarely look like clean reference architectures. A configuration may differ from a common baseline because there is a legitimate workload requirement behind it. A database may use a recovery model that would not be my default choice, but changing it without understanding the surrounding backup and recovery process would be irresponsible.</p>



<p class="wp-block-paragraph">The same applies to older applications, service accounts, maintenance jobs or settings that teams have carried forward through several generations of infrastructure. The first question is therefore not whether something matches a checklist. I want to know why it is configured that way and whether the reason still applies.</p>



<p class="wp-block-paragraph">Sometimes the answer is straightforward. The configuration is intentional, the team understands its impact and there is no reason to change it. In other cases nobody can explain where it came from. That does not make the setting wrong either, but it changes the situation because an unexplained dependency makes it harder to predict what will happen during maintenance, migration or failure.</p>



<h2 class="wp-block-heading">Successful jobs can still hide unanswered questions</h2>



<p class="wp-block-paragraph">SQL Agent jobs are a good example because they are easy to classify as healthy. A job runs successfully every night, there are no failed steps and monitoring stays green. Technically, SQL Server Agent has done what it was told to do. Microsoft describes <a href="https://learn.microsoft.com/en-us/ssms/agent/sql-server-agent">SQL Server Agent</a> as the service responsible for executing scheduled administrative tasks. The more useful operational question is whether the job is still doing what the organization believes it is doing.</p>



<p class="wp-block-paragraph">Someone may have created a maintenance routine when the database was much smaller or the workload looked very different. A data-processing job may still run after the surrounding application has changed. Backup jobs may successfully create backup files even though nobody has tested the complete recovery process for quite some time. Microsoft explicitly recommends testing a backup and restore strategy by actually restoring backups and recovering the database. Its <a href="https://learn.microsoft.com/en-us/sql/relational-databases/backup-restore/back-up-and-restore-of-sql-server-databases?view=sql-server-ver17">SQL Server backup and restore guidance</a> provides the technical background. The job history alone cannot answer those questions.</p>



<p class="wp-block-paragraph">The same principle applies elsewhere. Monitoring can remain green because the configured thresholds are being met, while the alerts themselves no longer represent the failures that actually matter to the application. Documentation can also exist without being particularly useful if it describes an architecture that has already changed several times. These are not necessarily urgent problems, but they are areas where the perceived state of the environment and the actual operational understanding may have drifted apart.</p>



<h2 class="wp-block-heading">Findings become useful when they have a consequence</h2>



<p class="wp-block-paragraph">A technical review can easily produce a long list of observations. That is not particularly difficult. The harder part is deciding which observations actually matter and what their operational consequence is.</p>



<figure class="wp-block-image size-large"><img fetchpriority="high" decoding="async" width="1024" height="576" src="https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-health-check-observation-context-consequence-priority-1024x576.png" alt="Infographic showing how a SQL Server Health Check turns a technical observation into context, operational consequence and priority." class="wp-image-21094" srcset="https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-health-check-observation-context-consequence-priority-1024x576.png 1024w, https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-health-check-observation-context-consequence-priority-300x169.png 300w, https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-health-check-observation-context-consequence-priority-768x432.png 768w, https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-health-check-observation-context-consequence-priority-1536x864.png 1536w, https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-health-check-observation-context-consequence-priority.png 1672w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">Suppose a backup configuration looks unusual. Before assigning a severity to it, I would want to understand whether it affects the ability to recover the system within the expected time. An undocumented configuration change is different. There, the relevant question may be what depends on it and what could happen if somebody changes it during an upgrade or troubleshooting session.</p>



<p class="wp-block-paragraph">If a SQL Server instance is part of a high-availability design, the interesting evidence is not simply that the configuration exists. I would want to know whether the team has tested the complete failover path recently enough to trust it, including the application side. The consequence gives the finding its weight.</p>



<p class="wp-block-paragraph">Without that context, a report can quickly become a collection of technical differences that all appear equally important. In reality, they rarely are. A minor configuration deviation with no practical effect may need nothing more than documentation. An untested recovery process deserves much more attention, even if every backup job has been successful for months. The same principle applies to lifecycle decisions: as I discussed in <a href="https://www.craftedsql.com/sql-server-2016-end-of-support-delayed-decisions/">SQL Server 2016 End of Support: The Real Problem Is Not the Date</a>, a system can continue to look stable while the operational risk around it changes. A security-related finding may require action for yet another reason.</p>



<h2 class="wp-block-heading">A SQL Server Health Check should help set priorities</h2>



<p class="wp-block-paragraph">I am not particularly interested in producing the longest possible findings list. A report containing fifty observations may look thorough, but it does not necessarily help the person responsible for the environment decide what to do on Monday morning.</p>



<p class="wp-block-paragraph">Some findings really do need prompt attention, while others are better addressed during the next planned maintenance window or migration project. In a few cases, documenting why something exists may be enough. There can also be situations where changing a technically unusual configuration would introduce more risk than leaving it alone.</p>



<p class="wp-block-paragraph">That distinction matters because every organization operates under constraints. Time is limited, maintenance windows are limited and technical changes compete with other work. A useful SQL Server Health Check should therefore make the difference between an interesting technical observation and an operational priority visible.</p>



<p class="wp-block-paragraph">The workload also changes the evaluation. An internal SQL Server used by a small application has different requirements from a platform supporting business-critical transactions around the clock. The same configuration may therefore deserve very different treatment in those two environments. There is no meaningful severity rating without understanding what the system is expected to do and what happens when it cannot.</p>



<h2 class="wp-block-heading">What I want to understand after the review</h2>



<p class="wp-block-paragraph">After reviewing an existing environment, I want the people responsible for it to have a clearer picture of how much of the current setup is intentional and understood. That includes the dependencies that matter during maintenance or failure. I also want to know whether backup and recovery processes still match the expectations placed on them, and whether scheduled operations are still doing what everyone assumes they are doing.</p>



<p class="wp-block-paragraph">Unusual configurations do not worry me simply because they are unusual. What matters is whether somebody can explain why they exist and what depends on them. From there, prioritization becomes much easier.</p>



<p class="wp-block-paragraph">Which findings could affect an upcoming upgrade? Which ones increase recovery risk? Where is the environment relying on knowledge that currently exists only in someone&#8217;s head? And which observations are interesting but realistically do not require any change? Those are much more useful questions than whether a server passes or fails a generic checklist.</p>



<p class="wp-block-paragraph">A SQL Server environment can run well today while still containing decisions that deserve another look. A good Health Check helps identify which of those decisions matter before an upgrade, migration or incident forces the team to find out under pressure.</p>



<p class="has-small-font-size wp-block-paragraph">Foto von <a href="https://unsplash.com/de/@tecnic?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">TECNIC Bioprocess Solutions</a> auf <a href="https://unsplash.com/de/fotos/ein-mann-misst-ein-metallbauteil-sorgfaltig-ab-EA7QL-GJWRM?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-health-check-operational-risk/">SQL Server Health Checks: Small Operational Gaps That Become Expensive Later</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>SQL Server Vector Search: What to Evaluate Before Production</title>
		<link>https://www.craftedsql.com/sql-server-vector-search-production-readiness/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 13 Aug 2026 09:28:00 +0000</pubDate>
				<category><![CDATA[Architecture & Setup]]></category>
		<category><![CDATA[Production Readiness]]></category>
		<category><![CDATA[SQL Server 2025]]></category>
		<category><![CDATA[Vector Search]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21060</guid>

					<description><![CDATA[<p>SQL Server 2025 adds native support for vector data and vector search. Applications can keep embeddings alongside relational data and use the SQL Database Engine to find semantically similar information. A small proof of concept is easy enough to build: store an embedding, calculate the distance between vectors and return similar records. The questions I ... <a title="SQL Server Vector Search: What to Evaluate Before Production" class="read-more" href="https://www.craftedsql.com/sql-server-vector-search-production-readiness/" aria-label="Read more about SQL Server Vector Search: What to Evaluate Before Production">Read more</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-vector-search-production-readiness/">SQL Server Vector Search: What to Evaluate Before Production</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">SQL Server 2025 adds native support for vector data and vector search. Applications can keep embeddings alongside relational data and use the SQL Database Engine to find semantically similar information.</p>



<p class="wp-block-paragraph">A small proof of concept is easy enough to build: store an embedding, calculate the distance between vectors and return similar records.</p>



<p class="wp-block-paragraph">The questions I find more relevant begin when vector data becomes part of a production workload. It then has to fit into the same operational environment as the rest of the database, including capacity planning, performance testing, security, recovery and ongoing support.</p>



<p class="wp-block-paragraph">At that stage, I want to know how the workload behaves in the environment where it will actually run. A working vector search alone does not answer that.</p>



<h2 class="wp-block-heading">SQL Server vector search can simplify an architecture</h2>



<p class="wp-block-paragraph">Native vector support allows relational data and vector data to stay in the same database platform.</p>



<p class="wp-block-paragraph">Microsoft describes the SQL Database Engine as an option for scenarios where applications need to search structured, unstructured and vector data together. If an application already relies heavily on SQL Server, keeping that workload in the same platform may avoid the need for a separate search service.</p>



<p class="wp-block-paragraph">The additional workload still has to run somewhere. In this design, it runs inside SQL Server, shares the same underlying resources and becomes part of the same operational responsibilities.</p>



<p class="wp-block-paragraph">Before choosing this architecture, I would want to understand what the vector workload adds to the existing database environment and whether the simplification at application level creates new pressure elsewhere.</p>



<p class="wp-block-paragraph"><a href="https://learn.microsoft.com/en-us/sql/sql-server/ai/vectors?view=sql-server-ver17">Microsoft: Vector search and vector indexes in the SQL Database Engine</a></p>



<h2 class="wp-block-heading">Vector data has a measurable database footprint</h2>



<p class="wp-block-paragraph">SQL Server stores vector values in an optimized binary format. With the default float32 representation, each dimension requires four bytes.</p>



<p class="wp-block-paragraph">A 1,536-dimensional vector uses 6,144 bytes, roughly 6 KB, for the vector values alone. The rest of the row and the surrounding database structures come on top of that.</p>



<figure class="wp-block-image size-large"><img decoding="async" width="1024" height="683" src="https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-vector-data-footprint-1536-dimensions-1024x683.png" alt="SQL Server vector data footprint showing that 1,536 float32 dimensions require about 6 KB per vector." class="wp-image-21064" srcset="https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-vector-data-footprint-1536-dimensions-1024x683.png 1024w, https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-vector-data-footprint-1536-dimensions-300x200.png 300w, https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-vector-data-footprint-1536-dimensions-768x512.png 768w, https://www.craftedsql.com/wp-content/uploads/2026/08/craftedsql-sql-server-vector-data-footprint-1536-dimensions.png 1536w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure>



<p class="wp-block-paragraph">In a small test database, that number is easy to overlook. With hundreds of thousands or millions of rows, vector data can become a noticeable part of the database footprint.</p>



<p class="wp-block-paragraph">Vector dimensions, row counts and expected growth therefore belong in normal capacity planning. A proof of concept with a few thousand rows tells us very little about how much space the same design may require later.</p>



<p class="wp-block-paragraph">The embedding model matters as well because different models use different vector dimensions. A choice made in the application or AI layer can have a direct effect on database size.</p>



<p class="wp-block-paragraph"><a href="https://learn.microsoft.com/en-us/sql/t-sql/data-types/vector-data-type?view=sql-server-ver17">Microsoft: Vector data type in SQL Server</a></p>



<h2 class="wp-block-heading">Vector search needs realistic production testing</h2>



<p class="wp-block-paragraph">Storage is only part of the workload. Search behavior also needs to be tested with realistic data volumes.</p>



<p class="wp-block-paragraph">Exact vector search calculates the distance between a query vector and the candidate vectors before returning the closest matches. As the number of candidates grows, SQL Server has more work to do.</p>



<p class="wp-block-paragraph">Microsoft gives fewer than 50,000 vectors as a general recommendation for exact search. A table itself can contain considerably more data if other predicates reduce the number of vectors involved in the actual search.</p>



<p class="wp-block-paragraph">The total table size therefore tells only part of the story. I also need to know how many vectors a particular query evaluates and how those searches behave alongside the existing workload.</p>



<p class="wp-block-paragraph">For a production-oriented test, I would measure execution time, CPU consumption and I/O with representative data volumes. I would then repeat the same tests while the rest of the application is active, because that is much closer to the environment the database will have to handle later.</p>



<p class="wp-block-paragraph">A fast search against a small dataset can be useful during development. It says little about how the same design will behave after months of data growth.</p>



<h2 class="wp-block-heading">Feature maturity becomes part of the architecture decision</h2>



<p class="wp-block-paragraph">For larger search sets, teams may eventually consider approximate search instead of evaluating every candidate vector.</p>



<p class="wp-block-paragraph">SQL Server 2025 includes approximate vector search and vector indexes for this scenario. Approximate search trades some exactness for better scalability and lower resource requirements.</p>



<p class="wp-block-paragraph">At the time of writing, approximate vector search and vector indexes are still Preview features in SQL Server 2025.</p>



<p class="wp-block-paragraph">I would not reject a feature simply because it carries a Preview label. I would, however, treat that dependency differently when evaluating a business-critical workload.</p>



<p class="wp-block-paragraph">The first question is whether the application needs that functionality at all. If it does, I also want to know what the fallback looks like if the planned implementation cannot be used as expected.</p>



<p class="wp-block-paragraph">The current vector implementation has other limitations that can affect an existing database design. Always Encrypted does not support the vector data type, and SQL Server does not allow vector columns as keys in traditional B-tree or columnstore indexes.</p>



<p class="wp-block-paragraph">For some applications, those limitations will have little practical effect. For others, they may conflict directly with an existing security or indexing design. That has to be checked against the actual environment.</p>



<h2 class="wp-block-heading">Production readiness is about the whole service</h2>



<p class="wp-block-paragraph">Once an application depends on embeddings and vector search, the operational model has to cover them as well.</p>



<p class="wp-block-paragraph">The security model still has to match the data the application stores. Access paths, permissions and data classification remain part of the same database design, even when some of the stored values are vectors.</p>



<p class="wp-block-paragraph">Recovery deserves the same attention. Restoring the database successfully is one technical step. The application also has to work with the restored data and provide the vector search functionality it depended on before the failure.</p>



<p class="wp-block-paragraph">Much of this is familiar DBA work. SQL Server 2025 introduces another type of workload, but sizing it, measuring its behavior and understanding its dependencies are familiar operational tasks.</p>



<p class="wp-block-paragraph">For a production decision, I would want evidence from the environment where the workload is expected to run: representative data volumes, realistic concurrent activity, understood feature dependencies and a recovery test that includes the application behavior. A successful demo would be the beginning of that evaluation, not the result.</p>



<p class="has-small-font-size wp-block-paragraph">Foto von <a href="https://unsplash.com/de/@tecnic?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">TECNIC Bioprocess Solutions</a> auf <a href="https://unsplash.com/de/fotos/ein-mann-misst-ein-metallbauteil-sorgfaltig-ab-EA7QL-GJWRM?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-vector-search-production-readiness/">SQL Server Vector Search: What to Evaluate Before Production</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>SQL Server Query Store: Why Performance History Matters</title>
		<link>https://www.craftedsql.com/sql-server-query-store-performance-history/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 06 Aug 2026 11:12:00 +0000</pubDate>
				<category><![CDATA[Performance & Optimization]]></category>
		<category><![CDATA[Execution Plans]]></category>
		<category><![CDATA[Performance Troubleshooting]]></category>
		<category><![CDATA[Query Store]]></category>
		<category><![CDATA[SQL Server Performance]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21056</guid>

					<description><![CDATA[<p>A SQL Server performance problem is not always still there when someone finally has time to investigate it. SQL Server Query Store becomes especially useful when the system has already returned to normal. Users may report that an application was slow around 10:00. An hour later, response times are back to normal. CPU usage looks ... <a title="SQL Server Query Store: Why Performance History Matters" class="read-more" href="https://www.craftedsql.com/sql-server-query-store-performance-history/" aria-label="Read more about SQL Server Query Store: Why Performance History Matters">Read more</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-query-store-performance-history/">SQL Server Query Store: Why Performance History Matters</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">A SQL Server performance problem is not always still there when someone finally has time to investigate it. SQL Server Query Store becomes especially useful when the system has already returned to normal.</p>



<p class="wp-block-paragraph">Users may report that an application was slow around 10:00. An hour later, response times are back to normal. CPU usage looks ordinary and there is no obvious blocking. The query mentioned by the application team now finishes without anything particularly suspicious in its execution plan.</p>



<p class="wp-block-paragraph">That is a common difficulty in performance troubleshooting: the system you are looking at now may no longer be the system users experienced during the incident.</p>



<p class="wp-block-paragraph">At that point, a current execution plan or a snapshot of server activity can still be useful, but it cannot tell the whole story. I want to know what happened around 10:00, not only what SQL Server looks like at 11:00.</p>



<p class="wp-block-paragraph">Query Store keeps <a href="https://learn.microsoft.com/en-us/sql/relational-databases/performance/monitoring-performance-by-using-the-query-store?view=sql-server-ver17" target="_blank" rel="noopener">historical information about queries, execution plans and runtime behavior</a>. That means an investigation does not have to rely entirely on the state of the database after the problem has disappeared.</p>



<h2 class="wp-block-heading">Start with when the problem happened</h2>



<p class="wp-block-paragraph">When I have a reasonably clear incident window, that is usually where I start.</p>



<p class="wp-block-paragraph">If the application team tells me that response times became poor between 09:45 and 10:15, I start with that period. I can then compare the affected queries with their behavior before or afterwards.</p>



<p class="wp-block-paragraph">I am not immediately looking for the most expensive query in the database.</p>



<p class="wp-block-paragraph">A query that consumes a lot of CPU every day may be completely unrelated to the incident. A query that normally gets little attention can suddenly become interesting. Its duration or execution pattern may have changed during exactly the period users reported a problem.</p>



<p class="wp-block-paragraph">The first comparison is therefore fairly simple.</p>



<p class="wp-block-paragraph">Did the affected query actually become slower around that time? Was another execution plan used? Did the query execute much more often than usual? And does the same behavior appear at other times as well?</p>



<p class="wp-block-paragraph">Those answers do not give me a diagnosis yet, but they help narrow the investigation.</p>



<h2 class="wp-block-heading">A different execution plan is a clue</h2>



<p class="wp-block-paragraph">One of the things Query Store makes easy to see is whether a query has used different execution plans over time.</p>



<p class="wp-block-paragraph">If a query usually runs with one plan and performed much worse while another plan was active, that deserves attention. It still does not mean that the second plan explains everything that happened.</p>



<p class="wp-block-paragraph">The workload may have been different. Parameter values may have changed. Data distribution may have influenced the optimizer. There may also have been resource pressure elsewhere on the server at the same time.</p>



<p class="wp-block-paragraph">This is why I am careful with the jump from “different plan” to “bad plan”.</p>



<p class="wp-block-paragraph">Sometimes the plan really is the immediate problem, and forcing a previously known plan can be a reasonable way to stabilize an important workload. During an incident, restoring acceptable service may be more important than understanding every detail immediately.</p>



<p class="wp-block-paragraph">Afterwards, however, I still want to know why the behavior changed in the first place.</p>



<p class="wp-block-paragraph">Otherwise, I have fixed the symptom without learning much from the incident.</p>



<h2 class="wp-block-heading">SQL Server Query Store puts execution plans into context</h2>



<p class="wp-block-paragraph">The execution plan is only one part of the picture.</p>



<p class="wp-block-paragraph">What I find more useful during a retrospective investigation is comparing the query across different periods. If duration increased sharply for half an hour and then returned to normal, I have a much narrower problem to investigate.</p>



<p class="wp-block-paragraph">The same applies when the pattern repeats.</p>



<p class="wp-block-paragraph">Perhaps the query becomes slower every morning around the same time. That immediately raises different questions. Is there a batch process running? Does the application generate a different workload during that period? Is data being imported? Is maintenance running somewhere in the environment?</p>



<p class="wp-block-paragraph">This becomes particularly relevant with ETL and analytical workloads, where growing processing windows can gradually begin to overlap with reporting or other activity. I look at that scenario in more detail in <a href="/sql-server-data-warehouse-performance/">SQL Server performance for data warehouse workloads</a>.</p>



<p class="wp-block-paragraph">The same pattern can appear across broader <a href="/sql-server-business-intelligence-performance/">SQL Server Business Intelligence workloads</a>, where ETL, semantic model processing and reporting may depend on the same database platform at different times of the day. Power BI is one specific example: a report or DirectQuery workload may perform normally most of the day but become slow when query patterns, concurrency or resource pressure change on the SQL Server side. The <a href="/sql-server-performance-power-bi/">SQL Server performance for Power BI workloads</a> page looks at that relationship in more detail.</p>



<p class="wp-block-paragraph">Query Store will not answer all of those questions.</p>



<p class="wp-block-paragraph">What it can do is show that the change happened and help define the period I need to investigate elsewhere.</p>



<p class="wp-block-paragraph">That becomes particularly useful when I compare the Query Store history with monitoring data from SQL Server, the operating system, storage or the virtualization platform. If several independent sources show a change during the same period, the investigation becomes much more focused.</p>



<p class="wp-block-paragraph">This also prevents one of the easier mistakes in troubleshooting: finding an interesting SQL Server observation and trying to make the entire incident fit around it. I discussed that problem in more detail in <a href="https://www.craftedsql.com/sql-server-bottleneck-triage-signals-causes/">SQL Server Bottleneck Triage: How to Separate Signals from Causes</a>.</p>



<h2 class="wp-block-heading">Historical data has to exist before you need it</h2>



<p class="wp-block-paragraph">There is a practical operational consequence to all of this.</p>



<p class="wp-block-paragraph">If Query Store is supposed to help with incidents that happened yesterday or during a short period overnight, it needs to retain useful data for long enough to make that possible.</p>



<p class="wp-block-paragraph">That sounds obvious, but it is something I would rather decide while the system is healthy than discover during an investigation.</p>



<p class="wp-block-paragraph">Capture settings, storage limits, retention and the workload itself influence how useful the available history will be. There is no single configuration that fits every database. This should therefore be an intentional operational decision rather than something left untouched indefinitely.</p>



<p class="wp-block-paragraph">For an important workload, I want to know roughly how far back I can investigate. I also want to know whether the data being collected is useful for the type of incidents that matter to that system.</p>



<p class="wp-block-paragraph">SQL Server Query Store does not replace monitoring or the rest of SQL Server troubleshooting. It gives me another perspective that becomes particularly useful once the immediate symptoms have disappeared.</p>



<p class="wp-block-paragraph">When someone tells me that the application was slow an hour ago, the first question is no longer only what SQL Server is doing now.</p>



<p class="wp-block-paragraph">I want to understand what was different when users actually noticed the problem.</p>
<p>The post <a href="https://www.craftedsql.com/sql-server-query-store-performance-history/">SQL Server Query Store: Why Performance History Matters</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>SQL Server Backup Strategy: Why Successful Backups Are Not Enough</title>
		<link>https://www.craftedsql.com/sql-server-backup-strategy/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 30 Jul 2026 07:32:00 +0000</pubDate>
				<category><![CDATA[DBA Essentials]]></category>
		<category><![CDATA[Disaster Recovery]]></category>
		<category><![CDATA[Restore Testing]]></category>
		<category><![CDATA[SQL Server Backup]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21047</guid>

					<description><![CDATA[<p>Most SQL Server environments have a backup process, but not every SQL Server backup strategy reflects actual recovery requirements. Teams may run jobs according to a defined schedule, monitor successful completion and retain the files for a specified period. Recovery readiness means that the organization can restore the required databases to an appropriate point in ... <a title="SQL Server Backup Strategy: Why Successful Backups Are Not Enough" class="read-more" href="https://www.craftedsql.com/sql-server-backup-strategy/" aria-label="Read more about SQL Server Backup Strategy: Why Successful Backups Are Not Enough">Read more</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-backup-strategy/">SQL Server Backup Strategy: Why Successful Backups Are Not Enough</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Most SQL Server environments have a backup process, but not every SQL Server backup strategy reflects actual recovery requirements. Teams may run jobs according to a defined schedule, monitor successful completion and retain the files for a specified period.</p>



<p class="wp-block-paragraph">Recovery readiness means that the organization can restore the required databases to an appropriate point in time, on suitable infrastructure and within a timeframe that supports the affected business process. The team also needs access to the necessary backup files, encryption keys, permissions, documentation and technical knowledge.</p>



<p class="wp-block-paragraph">A successful backup is therefore a prerequisite for recovery. It does not prove that recovery will work.</p>



<h2 class="wp-block-heading">Start the SQL Server backup strategy with recovery requirements</h2>



<p class="wp-block-paragraph">A backup strategy should not begin with a job schedule. It should begin with the systems and business processes that the organization needs to recover.</p>



<p class="wp-block-paragraph">Different databases can have very different requirements. A reporting database that another source can recreate does not necessarily need the same protection as a transactional database supporting orders, production, healthcare processes or financial operations.</p>



<p class="wp-block-paragraph">Before choosing backup types and frequencies, the organization should answer a few basic questions:</p>



<ul class="wp-block-list">
<li>How much data loss would be acceptable?</li>



<li>How quickly must the database become available again?</li>



<li>Does the business require point-in-time recovery?</li>



<li>Which databases and applications depend on each other?</li>



<li>How long must older recovery points remain available?</li>
</ul>



<p class="wp-block-paragraph">These answers influence backup frequency, recovery model, retention, storage capacity and the infrastructure needed for a restore.</p>



<p class="wp-block-paragraph">Without that connection, backup plans often grow historically. Full backups continue because they have always run. Transaction log backups follow an interval that someone selected years ago. Retention periods reflect available storage rather than the time in which a data problem might remain undetected.</p>



<p class="wp-block-paragraph">The jobs may still complete successfully, but the strategy behind them remains unclear.</p>



<h2 class="wp-block-heading">Protect the complete recovery scope</h2>



<p class="wp-block-paragraph">User databases represent only one part of a SQL Server environment.</p>



<p class="wp-block-paragraph">Depending on the recovery scenario, the team may also need system databases, SQL Server logins, SQL Server Agent jobs, credentials, proxies, linked servers, certificates, encryption keys and application-specific configuration.</p>



<p class="wp-block-paragraph">The required scope depends on what has failed.</p>



<p class="wp-block-paragraph">Restoring an accidentally deleted object to a separate database differs from recovering a complete SQL Server instance after losing the original server. Recovering one independent database also differs from restoring several databases that must represent a consistent point in time.</p>



<p class="wp-block-paragraph">This distinction matters because user database backups alone may not rebuild the service that depends on them. A database can be online while the application remains unusable because logins, jobs, keys or external dependencies are missing.</p>



<p class="wp-block-paragraph">Teams should therefore define the backup scope around realistic recovery scenarios rather than only the database files that currently exist on the server.</p>



<h2 class="wp-block-heading">Design the SQL Server backup strategy deliberately</h2>



<p class="wp-block-paragraph">SQL Server provides full, differential and transaction log backups, and each type serves a different purpose. The appropriate combination depends on database size, acceptable data loss, required restore time and the operational effort that the team can support.</p>



<p class="wp-block-paragraph">Databases with a small acceptable data-loss window may require frequent transaction log backups, while larger databases may benefit from differential backups that reduce the amount of data processed during regular backup operations.</p>



<p class="wp-block-paragraph">Using every available backup type is not the objective. Teams need to understand what each backup contributes to the recovery chain and whether the complete sequence supports the required recovery point.</p>



<p class="wp-block-paragraph">The database recovery model must match this design. Under the FULL recovery model, regular transaction log backups maintain the log backup chain and support point-in-time recovery. The SIMPLE recovery model follows a different recovery approach and does not offer the same restore options.</p>



<p class="wp-block-paragraph">Microsoft provides a detailed overview of the available <a href="https://learn.microsoft.com/en-us/sql/relational-databases/backup-restore/recovery-models-sql-server?view=sql-server-ver17" target="_blank" rel="noopener">SQL Server recovery models and their effect on restore operations</a>.</p>



<h2 class="wp-block-heading">Retention must preserve usable recovery points</h2>



<p class="wp-block-paragraph">Retention belongs to the recovery design. It is more than a storage-cleanup setting.</p>



<p class="wp-block-paragraph">Teams do not detect every data problem immediately. Incorrect changes, logical corruption, application defects or administrative mistakes can remain unnoticed for days or weeks. If all usable recovery points from before the event have already expired, a technically successful backup process may still fail to provide the required data.</p>



<p class="wp-block-paragraph">Retention also needs to preserve complete recovery chains. A transaction log backup has little value on its own when the required full backup or another necessary part of the chain is no longer available.</p>



<p class="wp-block-paragraph">The appropriate retention period depends on how long an issue might remain undetected, which historical recovery points the organization needs and how quickly database growth increases storage demand.</p>



<h2 class="wp-block-heading">Backup storage must survive the relevant failure</h2>



<p class="wp-block-paragraph">The backup destination requires the same attention as the backup schedule.</p>



<p class="wp-block-paragraph">When production data and backups share the same server, storage platform or administrative boundary, one incident can affect both. Storage failures, ransomware, access problems and administrative mistakes may remove the production system and its recovery path at the same time.</p>



<p class="wp-block-paragraph">A practical design may therefore use several copies for different purposes: one copy for fast restores, one protected or immutable copy and another copy outside the primary location.</p>



<p class="wp-block-paragraph">The fastest backup destination is not necessarily the most resilient recovery source. Teams need to balance backup performance, protection and restore availability.</p>



<h2 class="wp-block-heading">Keep encryption keys and credentials recoverable</h2>



<p class="wp-block-paragraph">Protecting backup files can create additional dependencies. The organization must retain everything required to access and decrypt those files later.</p>



<p class="wp-block-paragraph">Encrypted SQL Server backups, encrypted databases and secured repositories can reduce exposure, but certificates, private keys, passwords and access procedures must remain protected and recoverable as well.</p>



<p class="wp-block-paragraph">A backup that nobody can decrypt is not a usable backup.</p>



<p class="wp-block-paragraph">The organization should store the required keys and credentials separately from the systems they protect. Teams should restrict access without making recovery dependent on one administrator, one undocumented password or one server that may be unavailable during an incident.</p>



<p class="wp-block-paragraph">This becomes especially relevant when the environment uses Transparent Data Encryption, SQL Server backup encryption, third-party backup products or separately secured repositories. Microsoft documents the available <a href="https://learn.microsoft.com/en-us/sql/relational-databases/backup-restore/backup-encryption?view=sql-server-ver17" target="_blank" rel="noopener">SQL Server backup encryption options and their prerequisites</a>.</p>



<p class="wp-block-paragraph">Security and recoverability should support each other. Strong protection without a workable recovery path creates a different operational risk.</p>



<h2 class="wp-block-heading">Monitor backup coverage, continuity and availability</h2>



<p class="wp-block-paragraph">A green backup job confirms that the operation reported success. It does not prove that every required database has a backup or that the complete recovery chain remains available.</p>



<p class="wp-block-paragraph">Useful monitoring should therefore cover more than the final job result. It should verify that all expected databases are included, transaction log backups run at the intended interval, storage and retention cleanup remain healthy, and alerts reach someone who can act on them.</p>



<p class="wp-block-paragraph">Missing backups require particular attention.</p>



<p class="wp-block-paragraph">A newly created database may never become part of an older job configuration. A renamed instance, changed storage path or failed cleanup task may gradually create a gap without producing an obvious database error. Monitoring that checks only for failed jobs may never detect the missing backup.</p>



<p class="wp-block-paragraph">Teams should regularly compare the databases that exist with the databases covered by the backup process. This simple administrative control can reveal gaps that job-based monitoring misses.</p>



<p class="wp-block-paragraph">Changes in duration also provide useful information. A backup that takes much longer than before may point to database growth, slower storage, network limitations or changes in compression and encryption. The longer duration does not automatically indicate a problem, but it may affect maintenance windows and the time required to move or restore the backup.</p>



<h2 class="wp-block-heading">Keep recovery documentation aligned with the environment</h2>



<p class="wp-block-paragraph">Teams often create recovery documentation once and then treat it as complete. SQL Server environments do not remain static.</p>



<p class="wp-block-paragraph">Databases grow, applications change, servers are replaced, storage paths move and credentials expire. New encryption or security requirements introduce additional dependencies. A restore procedure that worked two years ago may no longer describe the current environment.</p>



<p class="wp-block-paragraph">Useful documentation should make the recovery prerequisites visible. It should include backup locations, retention periods, expected backup frequency, recovery models, encryption requirements, dependencies, compatible recovery infrastructure and the validation steps that follow a restore.</p>



<p class="wp-block-paragraph">Someone other than the administrator who originally configured the jobs should also be able to use the documentation.</p>



<p class="wp-block-paragraph">Complex incidents will still require technical judgement, and no checklist can replace that. Good documentation should nevertheless prevent teams from spending critical recovery time rediscovering basic facts about their own systems.</p>



<h2 class="wp-block-heading">Restore testing must influence your SQL Server backup strategy</h2>



<p class="wp-block-paragraph">Teams sometimes treat restore testing as an occasional technical exercise. It should instead form part of the <strong>SQL Server backup strategy</strong>.</p>



<p class="wp-block-paragraph">A restore test answers questions that backup monitoring cannot answer:</p>



<ul class="wp-block-list">
<li>Can SQL Server read the backup files?</li>



<li>Does the complete recovery chain exist?</li>



<li>Can the team access the required encryption keys and credentials?</li>



<li>Does suitable recovery infrastructure exist?</li>



<li>Can the restore reach the required recovery point?</li>



<li>How long does the restore actually take?</li>



<li>Does the documented procedure still match the environment?</li>
</ul>



<p class="wp-block-paragraph">The result should not end with a note that the test succeeded or failed. It should feed back into the backup design.</p>



<p class="wp-block-paragraph">If the restore takes longer than the agreed recovery time, the organization may need to change its backup sequence, storage layout, network capacity or recovery infrastructure.</p>



<p class="wp-block-paragraph">Difficult access to the required keys points to a key-management problem. Insufficient free storage requires an adjustment to capacity planning. Missing dependencies in the documentation should become part of the recovery scope.</p>



<p class="wp-block-paragraph">Restore testing does more than validate existing backups. It tests the assumptions behind the strategy.</p>



<p class="wp-block-paragraph">Not every test needs to simulate the loss of an entire data center. The scenario should, however, reflect situations the organization may realistically face. Examples include recovering a database after accidental data loss, performing point-in-time recovery or restoring a large encrypted database to alternate infrastructure.</p>



<p class="wp-block-paragraph">Bringing a database to the ONLINE state confirms that SQL Server could process the backup files. Depending on the test objective, the team may still need to validate application access, dependencies and data before declaring the recovery successful.</p>



<p class="wp-block-paragraph">Microsoft maintains an overview of <a href="https://learn.microsoft.com/en-us/sql/relational-databases/backup-restore/back-up-and-restore-of-sql-server-databases?view=sql-server-ver17" target="_blank" rel="noopener">backup and restore strategies for SQL Server databases</a>, including the relationship between backup types, recovery models and restore operations.</p>



<h2 class="wp-block-heading">Review your backup and recovery strategy after changes</h2>



<p class="wp-block-paragraph">A backup strategy that worked for a 200 GB database may no longer meet the same recovery target after the database grows to several terabytes.</p>



<p class="wp-block-paragraph">A restore that once completed within one hour may take considerably longer after storage changes, encryption changes or continued data growth. A migration may invalidate paths, permissions and documented procedures. New backup software may introduce different credentials, repositories and operational dependencies.</p>



<p class="wp-block-paragraph">Teams should therefore review the <strong>SQL Server backup strategy</strong> after relevant changes, including:</p>



<ul class="wp-block-list">
<li>significant database growth</li>



<li>SQL Server upgrades or migrations</li>



<li>changes to storage or backup software</li>



<li>new encryption and security requirements</li>



<li>application or dependency changes</li>



<li>revised recovery objectives</li>
</ul>



<p class="wp-block-paragraph">A review does not have to redesign everything. It should confirm that the original assumptions, available infrastructure and measured restore times still match the current environment.</p>



<h2 class="wp-block-heading">When an independent backup and recovery review is useful</h2>



<p class="wp-block-paragraph">An independent review can help when backup jobs run successfully but the organization remains uncertain about its recovery capability.</p>



<p class="wp-block-paragraph">Typical indicators include restore tests that have not taken place recently, unknown recovery times, historically grown job configurations, several backup tools or storage locations, unclear encryption dependencies and incomplete documentation.</p>



<p class="wp-block-paragraph">Such a review should not replace the internal team or introduce a more complicated backup product. Its purpose is to verify that the existing strategy supports the required recovery scenarios, identify practical gaps and define reasonable next steps.</p>



<p class="wp-block-paragraph">A focused SQL Server review can also help distinguish between a technical configuration issue and a broader gap in recovery planning. The <a href="/sql-server-collaboration/">Working Together</a> page explains how an initial review and collaboration can be structured.</p>



<h2 class="wp-block-heading">Recovery readiness begins before the incident</h2>



<p class="wp-block-paragraph">Backups are essential, but their value becomes visible only when teams can use them under real recovery conditions.</p>



<p class="wp-block-paragraph">A reliable <strong>SQL Server backup strategy</strong> connects recovery requirements with backup types, retention, storage protection, key management, monitoring, documentation and restore testing.</p>



<p class="wp-block-paragraph">The number of successful jobs is not the strongest indicator of recovery readiness. The organization should be able to demonstrate that it can restore the required data within an acceptable timeframe and on infrastructure that remains available during the relevant failure scenario.</p>



<p class="wp-block-paragraph">That creates the technical and administrative foundation for recovery.</p>



<p class="wp-block-paragraph">The next article in this two-part series will look at what happens when the organization has to use that foundation: who leads a restore, who makes the necessary decisions, how the teams validate the result and who approves the return to service.</p>



<p class="has-small-font-size wp-block-paragraph">Foto von <a href="https://unsplash.com/de/@hdbernd?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Bernd 📷 Dittrich</a> auf <a href="https://unsplash.com/de/fotos/eine-schwarz-weisse-treppe-QMx2Ll4YFuA?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-backup-strategy/">SQL Server Backup Strategy: Why Successful Backups Are Not Enough</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>SQL Server Blocking: What the Head Blocker Does Not Tell You</title>
		<link>https://www.craftedsql.com/sql-server-blocking-head-blocker/</link>
		
		<dc:creator><![CDATA[Bjoern@CraftedSQL]]></dc:creator>
		<pubDate>Thu, 23 Jul 2026 11:21:00 +0000</pubDate>
				<category><![CDATA[SQL Server Problems]]></category>
		<category><![CDATA[Blocking]]></category>
		<category><![CDATA[Head Blocker]]></category>
		<category><![CDATA[Root Cause Analysis]]></category>
		<category><![CDATA[SQL Server Troubleshooting]]></category>
		<category><![CDATA[Transactions]]></category>
		<guid isPermaLink="false">https://www.craftedsql.com/?p=21040</guid>

					<description><![CDATA[<p>A SQL Server blocking chain shows which session is holding up other work. It does not explain why the transaction remained open long enough to affect production.</p>
<p>The post <a href="https://www.craftedsql.com/sql-server-blocking-head-blocker/">SQL Server Blocking: What the Head Blocker Does Not Tell You</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">When users report that an application has stopped responding, a SQL Server blocking chain can provide an immediate explanation.</p>



<p class="wp-block-paragraph">One session is waiting for another. More sessions may be waiting behind it. At the top of that chain is the head blocker.</p>



<p class="wp-block-paragraph">Finding that session matters. During an active incident, it helps assess which workload is affecting the system and whether intervention may be necessary.</p>



<p class="wp-block-paragraph">It does not yet explain why normal transaction processing developed into a production problem.</p>



<h2 class="wp-block-heading">SQL Server blocking does not automatically mean something is broken</h2>



<p class="wp-block-paragraph">SQL Server coordinates concurrent access through locks and, depending on the isolation model, row versioning. Blocking occurs when one session holds a lock and another session requests an incompatible lock on the same resource. Microsoft provides a detailed technical guide to <a href="https://learn.microsoft.com/en-us/troubleshoot/sql/database-engine/performance/understand-resolve-blocking">understanding and resolving SQL Server blocking</a>.</p>



<p class="wp-block-paragraph">That is normal transaction behaviour. A database system must protect data consistency while work is in progress.</p>



<p class="wp-block-paragraph">The operational question is not whether blocking exists. It is whether sessions wait long enough to affect users, scheduled processing, batch completion, or dependent workloads.</p>



<p class="wp-block-paragraph">A brief wait between two transactions may be harmless. The same locking relationship becomes a problem when the transaction remains open, the number of waiting sessions grows, or a time-sensitive process cannot continue. This wider relationship between <a href="https://www.craftedsql.com/sql-server-blocking-concurrency-enough-cpu/">SQL Server blocking and concurrency</a> is easy to miss when infrastructure metrics still appear normal.</p>



<p class="wp-block-paragraph">The blocking chain describes the current dependency. It does not explain why that dependency lasted long enough to matter.</p>



<h2 class="wp-block-heading">What the head blocker can and cannot tell you</h2>



<p class="wp-block-paragraph">The head blocker is the session at the top of the active blocking chain. Other requests are waiting directly or indirectly for resources held by that session.</p>



<p class="wp-block-paragraph">During an incident, I treat that as an operational fact. I want to understand what the session is doing, how many requests are affected, how long they have been waiting, and whether the impact is still growing. Current request information, including the blocking session and wait details, is available through <a href="https://learn.microsoft.com/en-us/sql/relational-databases/system-dynamic-management-objects/sys-dm-exec-requests-transact-sql?view=sql-server-ver17"><code>sys.dm_exec_requests</code></a>.</p>



<p class="wp-block-paragraph">That information helps me decide whether I need to protect a critical workload or restore service.</p>



<p class="wp-block-paragraph">I would not treat the head blocker as the root cause based on that information alone.</p>



<p class="wp-block-paragraph">The session may be running an expected application transaction. It may be idle while an earlier statement has left a transaction open. The work may take longer than usual because more data is involved, the execution plan has changed, or another process now overlaps with it.</p>



<p class="wp-block-paragraph">The blocking chain does not distinguish between those situations. It shows where the waiting begins, not what caused the transaction to hold locks for that length of time.</p>



<h2 class="wp-block-heading">Transaction duration and timing often matter more than the session ID</h2>



<p class="wp-block-paragraph">A session does not have to execute incorrect SQL to cause serious blocking.</p>



<p class="wp-block-paragraph">An application can begin a transaction, perform database work, and then wait for additional processing before committing. A large update may be legitimate but hold locks longer than the surrounding workload can tolerate. A query may read or modify more data than expected because its access path has changed.</p>



<p class="wp-block-paragraph">Timing can be just as important.</p>



<p class="wp-block-paragraph">Two processes may run without conflict for months when they execute at different times. A schedule change, increased data volume, a slower upstream process, or a longer-running batch can make them overlap.</p>



<p class="wp-block-paragraph">Reporting workloads can become part of the same concurrency problem. For example, <a href="/sql-server-performance-power-bi/">Power BI workloads using SQL Server</a> may compete with transactional or batch activity when queries overlap, potentially making the impact visible first as slow reports rather than as an obvious database problem.</p>



<p class="wp-block-paragraph">Neither process has to be new. The workload relationship has changed.</p>



<p class="wp-block-paragraph">That is why I would not change an index, isolation level, application query, or SQL Server setting based only on a snapshot of the blocking chain. The same principle applies throughout <a href="https://www.craftedsql.com/sql-server-bottleneck-triage/">SQL Server bottleneck triage</a>: the first visible signal needs to be tested against workload and runtime context before it becomes the basis for a change.</p>



<p class="wp-block-paragraph">Before changing the system, the investigation needs to explain what kept the transaction open and why competing work reached the same resources at that moment.</p>



<h2 class="wp-block-heading">Historical context matters when the blocking has already ended</h2>



<p class="wp-block-paragraph">A reliable SQL Server blocking analysis needs more than the current session ID. Current session and transaction data can describe what is happening now. If the issue has already ended or only appears at certain times, historical evidence becomes more important. As discussed in <a href="https://www.craftedsql.com/sql-server-monitoring-is-not-diagnosis/">SQL Server Monitoring Is Not Diagnosis</a>, recorded data creates visibility, but the interpretation still depends on workload and operational context.</p>



<p class="wp-block-paragraph">Extended Events or monitoring data may preserve the blocking relationship itself. Application logs, job history, and deployment timelines can help reconstruct the surrounding events. Query Store can add execution-plan and runtime context, but it does not record the blocking chain.</p>



<p class="wp-block-paragraph">The relevant evidence depends on the situation. The investigation may need to determine whether a transaction remained open while the application waited for another service, whether a batch started earlier than usual, or whether a plan or data-volume change extended the runtime enough to create an overlap.</p>



<p class="wp-block-paragraph">Not every incident requires every data source. It requires enough evidence to explain the duration and timing before a permanent change is made.</p>



<h2 class="wp-block-heading">Incident response and root-cause analysis are different tasks</h2>



<p class="wp-block-paragraph">Production pressure often forces a decision before the full cause is known.</p>



<p class="wp-block-paragraph">Terminating the head blocker may restore service. It also aborts the transaction and starts a rollback. Depending on the amount of work already completed, that rollback can take time and may continue to hold resources while SQL Server reverses the changes.</p>



<p class="wp-block-paragraph">Leaving the session running avoids that immediate rollback but may extend the impact on waiting processes.</p>



<p class="wp-block-paragraph">There is no universal answer. The decision depends on the importance of the blocking transaction, the number and criticality of affected sessions, whether the work is still progressing, and the likely operational effect of a rollback.</p>



<p class="wp-block-paragraph">That is incident response.</p>



<p class="wp-block-paragraph">Root-cause analysis begins once the immediate pressure is under control. It asks what should change so that the same workload pattern does not create the same impact again.</p>



<p class="wp-block-paragraph">The distinction matters. An emergency termination does not prove that the session or query was badly designed. Restoring service also does not mean that the underlying problem has been resolved.</p>



<h2 class="wp-block-heading">What should be reviewed before changing the system</h2>



<p class="wp-block-paragraph">The evidence may point to different parts of the environment.</p>



<p class="wp-block-paragraph">A transaction may need to be shorter or split into smaller units. A query may need a better access path. Two scheduled processes may need different execution windows. An application may hold a transaction open while it performs work that belongs outside the transaction.</p>



<p class="wp-block-paragraph">In other cases, the locking behaviour is expected but the isolation model no longer fits the workload. That is an architectural decision and should not be made during an incident.</p>



<p class="wp-block-paragraph">The most useful correction may not be a database change at all. The application workflow, batch design, scheduling, or monitoring may need adjustment.</p>



<p class="wp-block-paragraph">Application timeouts may also require review when they do not match the expected processing time. Increasing a timeout does not remove the blocking and cannot replace an understanding of the underlying workload.</p>



<p class="wp-block-paragraph">The head blocker remains important because it shows where the active wait begins. It may provide enough information to decide whether immediate action is required.</p>



<p class="wp-block-paragraph">It is not enough to decide what should be changed permanently.</p>



<p class="wp-block-paragraph">For that, the transaction duration, query behaviour, workload overlap, and timeline still have to explain why ordinary locking became a production issue.</p>



<p class="has-small-font-size wp-block-paragraph">Foto von <a href="https://unsplash.com/de/@s_tsuchiya?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Se. Tsuchiya</a> auf <a href="https://unsplash.com/de/fotos/eine-treppe-in-einem-gebaude-BUOSFehUBmg?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a></p>
<p>The post <a href="https://www.craftedsql.com/sql-server-blocking-head-blocker/">SQL Server Blocking: What the Head Blocker Does Not Tell You</a> appeared first on <a href="https://www.craftedsql.com">CraftedSQL</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
