Azure AD Connect 2.0 Won’t Start - 夜莺博客

Azure AD Connect 2.0 Won’t Start

原文:Azure AD Connect 2.0 Won’t Start — theDXT (Daniel Keer)

I recently ran into an issue where an install of Azure AD Connect failed to start. It seems like the root cause was due to the SQLLocalDB Model database becoming corrupt, which caused it to fail at upgrading itself. This is a known issue in versions older than 2.1.1.0 of Azure AD Connect.

While looking at the event logs it looks like the chain of events was that it tried to do the auto upgrade as auto upgrade is enabled but then failed to restart the SQLLocalDB due to the corruption which then caused Azure AD Connect to break.

Image 1
The application event log is filled with errors with SQLLocalDB failing to start.

The exact error I was getting was Event ID 528 with the following Windows API call WaitForMultipleObjects returned error code: 575. Windows system error message is: {Application Error}The application was unable to start correctly (0x%lx). Click OK to close the application.Reported at line: 3714.

Fortunately this setup of Azure AD Connect was only used for syncing users and password hash sync so the impact was minimal and there was a second Azure AD Connect agent that could’ve been failed over to if needed.

In order to fix this you can’t just install a new version of Azure AD Connect on top of it. You have to actually fix the corruption.

Thankfully Microsoft has an article on how to fix it. You can find that here https://docs.microsoft.com/en-us/troubleshoot/azure/active-directory/resolve-model-database-corruption-sqllocaldb#mitigation

Following those steps I was able to fix the issue and get Azure AD Connect started and running again.

Image 2
I then waited for the Azure AD Connect alerts to clear and it auto upgraded itself without issue this time.

It is worth understanding why this failure mode exists at all, because the version numbers matter and the same class of problem still turns up in the field on servers that have been left to look after themselves for a few years. The rest of this post covers the mechanics, the full mitigation, how to verify the repair, and how to stop it from happening again.

Why Azure AD Connect Uses SQL LocalDB at All

Microsoft Entra Connect — still widely known by its former name, Azure AD Connect — keeps its staging, connector and metaverse data in a SQL Server database. On a small or medium deployment the installer uses SQL Server Express LocalDB rather than a full SQL Server instance, which keeps the footprint tiny and removes the need for a separate database licence and maintenance story. LocalDB runs as a user-mode process, started on demand by the sync engine rather than as a permanent Windows service, and it stores its files under the profile of the service account.

That design has a consequence that is easy to overlook: LocalDB instances are created from a set of template databases, and one of those templates — the Model database — is used every single time a new LocalDB instance, database or connection is created. If the Model database becomes corrupt, nothing that needs a new database can start. LocalDB itself refuses to start, and the sync engine that depends on it fails with errors that point nowhere near the real cause.

The corruption itself usually arrives in one of three ways. An unclean shutdown of the host, a storage layer that acknowledged a write it did not actually commit, or — most commonly in the field — an in-place upgrade of Azure AD Connect that terminates and restarts LocalDB while the Model template is mid-write. That last pattern is exactly the one described in the original incident: the product auto-upgraded itself, failed to restart LocalDB, and left the tenant with a broken synchronisation agent that would not come back up on its own.

Symptoms to Look For

The failure is rarely announced clearly. The most reliable indicators are the following:

  • The Synchronization Service Manager will not open, or opens and reports that it cannot connect to the local database.
  • The Azure AD Connect Health agent shows the server as unhealthy, or the Connect Health portal reports no heartbeat from the sync server.
  • The Application event log fills with LocalDB errors. Event ID 528 carrying the WaitForMultipleObjects returned error code: 575 message is the signature failure, and the “application was unable to start correctly” phrasing is the tell that the underlying instance never reached a usable state.
  • Object changes stop flowing. Users created on-premises do not appear in the cloud, password changes do not propagate, and the delta sync cycle never completes — but users can still sign in with previously synced credentials, which makes the outage easy to miss.
  • A scheduled upgrade never completes. If auto upgrade is enabled and the service claims it is upgrading for an unusually long period, the underlying cause is frequently this corruption.

Because password hash synchronisation keeps working for existing credentials in many configurations, the most dangerous version of this incident is the silent one: the directory looks healthy until someone changes a password or creates a user, at which point the lag becomes visible in the worst possible way.

Before You Start

Fix the corruption on a server where you can afford to have the sync engine down for twenty minutes. The relevant preparation is:

  • Confirm the version. Run Get-ADSyncGlobalSettings or check the entry in Programs and Features. Anything older than 2.1.1.0 carries the known defect; anything 2.1.1.0 or newer should already handle the upgrade path correctly, but the Model database can still be corrupted by a bad shutdown, so the repair is not version-specific.
  • Take a snapshot or backup. If the server is a virtual machine, snapshot it. If it is physical, at minimum capture a copy of the %LOCALAPPDATA%\Microsoft\Microsoft SQL Server Local DB\Instances\ folder tree and record the configured sync settings. The mitigation deletes database files, and while nothing in your Entra directory is lost — that data lives in the cloud and in AD — the connector configuration is worth being able to restore.
  • Have a failover path. If a second staging-mode server exists, confirm that it is healthy and that you know how to promote it. Staging mode is the single most valuable protection in any Entra Connect deployment.
  • Sign in with the right rights. You need local administrator on the server and the credentials of the service account that owns the LocalDB instance, because the database files live in that profile.

The Mitigation

The full procedure Microsoft documents in the article linked above boils down to removing the corrupted databases and forcing LocalDB to rebuild them from a clean template, then letting Azure AD Connect recreate its own data. The sequence is:

  1. Stop the sync service. Run Stop-Service ADSync from an elevated PowerShell prompt, and confirm the process has actually exited — the service occasionally takes a while to shut down while it waits on a database call that will never return.
  2. Stop every LocalDB instance associated with the server. The supported approach is sqllocaldb stop MSSQLLocalDB, repeated for each instance name you see in sqllocaldb info. If a stop hangs, killing the stray sqlservr.exe process attached to the instance is acceptable at this point; you are about to delete its data anyway.
  3. Delete the affected instances and their files. Run sqllocaldb delete <instance> and then remove the leftover files under the service account profile so that no half-written Model database survives.
  4. Recreate the default instance and start it, using sqllocaldb create MSSQLLocalDB followed by sqllocaldb start MSSQLLocalDB. LocalDB regenerates the Model database from its installed template at this step — this is the actual repair.
  5. Restart the sync service with Start-Service ADSync, then open the Synchronization Service Manager and confirm that it connects without an error.

The most common mistake is skipping step three. Deleting the instance without removing the files, or removing some files but not the Model database files, leaves the corrupt template in place and the problem returns within a day or two. Be thorough.

Verifying the Repair

The service starting is necessary but not sufficient. Work through the following checks before you call the incident closed:

  • Run a manual delta synchronisation cycle and confirm it completes without errors, then run a full import and a full synchronisation if time allows. The Synchronization Service Manager reports the run profile results per connector.
  • Confirm the connector operations show plausible add, update and delete counts rather than zero across the board — zeros usually mean the connector is running but not actually reading from the directory.
  • Change a test user’s password on-premises and watch it propagate. This is the definitive end-to-end proof that password hash synchronisation is genuinely working rather than merely appearing healthy.
  • Check the Connect Health portal and wait for the alerts generated during the outage to clear. It can take a sync cycle or two, and clearing them is a clean marker that the agent is fully back.

Once the server was back and healthy in the original incident, it went on to auto upgrade itself without any trouble on the next attempt — a good sign that the corruption, and not the upgrade mechanism, was the real fault.

Preventing a Repeat

Rebuilding the Model database fixes the immediate problem but does nothing about the conditions that produced it. A few inexpensive habits make the difference:

  • Keep the agent current. Anything below 2.1.1.0 should be upgraded deliberately, in a maintenance window, rather than left to auto upgrade at an hour of Microsoft’s choosing.
  • Fix the shutdown path. Ensure the host is not hard-powered-off, and make sure any storage layer underneath it flushes writes before acknowledging them. Unclean shutdowns are the leading cause of this specific corruption.
  • Run a second server in staging mode. In the original incident the impact was minimal precisely because a second agent existed. Staging mode costs one licence and turns a directory-wide outage into a ten-minute failover.
  • Alert on sync staleness, not just on failure. A monitor that fires when the last successful sync is older than ninety minutes catches the silent variant of this failure long before a user notices.
  • Track the version and the sync health together. Entra Connect Health has improved considerably over the product’s life and will surface most of what you need, but pairing it with a simple scheduled check of the local service gives you coverage when the Health agent itself is the thing that is broken.

Frequently Asked Questions

Can I just reinstall Azure AD Connect over the top? No. That is the trap. The installer runs, appears to succeed, and then fails at exactly the same point, because the corrupt Model database is still there. Repair first, then upgrade.

Will I lose my sync configuration? Not normally — the connector and metaverse data live in the same LocalDB instance you are deleting, so you should expect to re-run the configuration wizard and, in the worst case, recreate custom sync rules. Export your custom rules beforehand if you have any, and use the documented procedure for backing up and restoring the configuration if you want the safer route.

Is this still a thing in the current release? The specific auto-upgrade defect is fixed in 2.1.1.0 and later, and the product has long since been renamed Microsoft Entra Connect. The Model database corruption itself is a SQL Server LocalDB behaviour, not an Entra Connect bug, so the mitigation remains relevant.

Should I move to a full SQL Server instead? For very large directories, or where you want normal SQL backup and monitoring, yes — that is the supported enterprise configuration. For a few tens of thousands of objects LocalDB is perfectly adequate, as long as you keep the platform patched and the server shut down cleanly.

Final Thoughts

This is one of those failures that looks catastrophic and turns out to be tractable, provided you resist the instinct to reinstall. Identify the Event ID 528 signature, work through the LocalDB mitigation properly — including deleting the files, not just the instances — verify with a real password change rather than a green service light, and then put the prevention measures in place so that the next upgrade you let run unattended does not undo the afternoon’s work.

If you have more than one directory to look after, it is also worth being able to work quickly at scale when the directory itself needs editing; the technique covered in mass editing ADSI values has saved me more than once during this kind of recovery. On the identity side, deciding early how Entra ID Conditional Access should behave protects the sign-in path while synchronisation is down, and if the tenant ever needs rebuilding from scratch, the process for a Microsoft 365 tenant deletion is worth reading before, not after.