GitHub

THB-P-SFTP-WIN Storage Troubleshooting and Cleanup

Find out what is filling the drives on THB-P-SFTP-WIN, the Health Align production SFTP server (Cerberus FTP Server), and free space safely.

Use this runbook when:

  • The low disk space - thb-p-sftp-win alert fires (C: or E: below 10% free)
  • C: free space is dropping faster than expected
  • A new meal vendor or plan is added to the SFTP copy job

The server at a glance

Azure VM THB-P-SFTP-WIN, resource group THB-PROD-SFTP-RG, East US 2, Health Align subscription. Not managed by Terraform.
Windows computer name HAWEB01. The VM was restored from a backup of the old HAWEB01, so logs and metrics use that name.
C: 512 GB OS disk. Holds C:\FTP, the local folders Cerberus serves to partners.
E: 300 GB data disk. Holds E:\archive (old logs, robocopy logs).
S: Azure file share \\hasftp.file.core.windows.net\ha-sftp-share, mapped globally at boot, so every account including SYSTEM sees it. The long-term copy of all SFTP data.

Scheduled tasks that matter for disk space:

Task Schedule What it does
\MoveFTPFiles Every 10 min, as HALocalAdmin Runs C:\FTP\moveFTP.ps1: robocopy moves partner uploads from C:\FTP to S:, and copies outgoing files from S: back to C:\FTP.
\ArchiveMealReferrals Daily 02:00, as SYSTEM Runs C:\FTP\archiveMeals.ps1: moves meal referral files older than 90 days from C:\FTP to S:\Archive\meals. Logs to E:\archive\meals-archive.log.
\SystemCleanup Daily 03:00, as HALocalAdmin Runs C:\DevOps\Powershell\Cleanup\cleanup.ps1: empties the recycle bin and temp folders.

IIS (W3SVC and WAS) is disabled on this server. Its sites were unreachable leftovers from HAWEB01; don’t re-enable it.

Access

There is no SSH shell: port 22 is Cerberus’s own SFTP listener. Use one of:

  • RDP for anything that writes (moves, deletes, task changes). Run PowerShell as administrator: UAC is on, and a normal window can copy files but fails every delete with ERROR 5 (Access is denied).
  • Azure Run Command for read-only checks without logging in. It runs as SYSTEM:
az vm run-command invoke -g THB-PROD-SFTP-RG -n THB-P-SFTP-WIN --command-id RunPowerShellScript --scripts "@check.ps1"

Run Command limits:

  • Only one command at a time per VM; a second one fails with Conflict until the first finishes.
  • Output is cut to the last ~4 KB. Print compact lines, not tables.
  • Pass the script as a file (@check.ps1). Inline scripts with quotes or & break in az.cmd on Windows.
  • Never list S: recursively through it. The share holds 2+ TB and the listing runs for 10+ minutes, blocking Run Command and slowing the share. To stop a stuck one, end the powershell -File scriptNN.ps1 child of RunCommandExtension.exe on the VM.

Triage

1. Check free space

Get-PSDrive C, E | Select-Object Name, @{n='UsedGB';e={[math]::Round($_.Used/1GB,2)}}, @{n='FreeGB';e={[math]::Round($_.Free/1GB,2)}}

2. Find what’s large on C:

Top-level folders first, then drill into the largest one by changing the path. This only reads C:, so it’s safe through Run Command.

Get-ChildItem C:\ -Directory -Force -ErrorAction SilentlyContinue | ForEach-Object {
  $gb = (Get-ChildItem $_.FullName -Recurse -File -Force -ErrorAction SilentlyContinue | Measure-Object Length -Sum).Sum / 1GB
  [pscustomobject]@{ GB = $gb; Path = $_.FullName }
} | Sort-Object GB -Descending | ForEach-Object { '{0,8:N2} GB  {1}' -f $_.GB, $_.Path }

3. Check the known growth sources

Check Expected Command
Meal referral files older than 90 days on C: 0, or one day’s batch before the 02:00 run See below
\ArchiveMealReferrals last result 0 Get-ScheduledTask -TaskName ArchiveMealReferrals | Get-ScheduledTaskInfo
\MoveFTPFiles last result 0 Get-ScheduledTask -TaskName MoveFTPFiles | Get-ScheduledTaskInfo
Stuck partial uploads (*.sftpcopy.tmp) None Get-ChildItem C:\FTP -Recurse -Filter *.sftpcopy.tmp
IIS W3SVC and WAS Stopped, Disabled Get-Service W3SVC, WAS

Meal referral files older than 90 days, outside incoming:

$cut = (Get-Date).AddDays(-90)
Get-ChildItem C:\FTP\*\*\referrals -Directory |
  Where-Object { $_.Parent.Name -in 'gafoods','icon','modify','momsmeals','homestyle' } |
  ForEach-Object { Get-ChildItem $_.FullName -Recurse -File -Force } |
  Where-Object { $_.FullName -notlike '*\incoming\*' -and $_.LastWriteTime -lt $cut } |
  Measure-Object Length -Sum | Select-Object Count, @{n='GB';e={[math]::Round($_.Sum/1GB,2)}}

Known causes

Cause Why it grows Fix
Meal referral exports in C:\FTP\<plan>\<vendor>\referrals Every export is a full snapshot (~23 a day per vendor folder). moveFTP.ps1 copies them from S: to C: and never deletes. Already handled: the copy-back lines use /maxage:90, and \ArchiveMealReferrals moves older files off. If old files pile up, check that task. If a new vendor or plan was added, see Adding a meal vendor or plan.
Stuck partial uploads (*.sftpcopy.tmp) A failed upload leaves a partial file, and moveFTP.ps1 skips *.tmp. Check that a complete copy (same name without .sftpcopy.tmp) reached S:, or that a later file replaced it, then delete. A partial file can never be processed.
IIS logs Only if IIS was re-enabled. Disable it again: Stop-Service WAS -Force; Set-Service W3SVC -StartupType Disabled; Set-Service WAS -StartupType Disabled
Anything new Use Moving a folder off C: once you know the data isn’t read in place.

Moving a folder off C:

Before moving anything, confirm nothing reads it from that path (Cerberus users’ folders, moveFTP.ps1, an app’s configured path). Move to S:\Archive\<name> for data that must be kept long term, or to E:\archive\<name> when it fits and is short-lived.

  1. Open PowerShell as administrator (the title bar must say Administrator).
  2. Check the destination has room (Get-PSDrive E / S: has ~3 TB).
  3. Dry run with /L, then run for real. Use /MT:32 for S: (many small files over the network; /MT:8 is ~5x slower) and keep the log on E::
robocopy <source> <dest> /S /E /DCOPY:T /COPY:DAT /MOV /NP /NFL /NDL /MT:32 /R:2 /W:5 /L
robocopy <source> <dest> /S /E /DCOPY:T /COPY:DAT /MOV /NP /NFL /NDL /MT:32 /R:2 /W:5 /LOG:E:\archive\<name>-move.log /TEE
  1. Don’t click inside the window while it runs. QuickEdit pauses the console, and robocopy freezes until you press Esc.
  2. Check the summary: FAILED must be 0. Robocopy reserves each file’s full size on the destination up front, so destination size isn’t a progress meter; the source shrinking is.
  3. If a file was copied but not deleted (for example the first run wasn’t elevated), a rerun skips it and leaves it on C:. Delete leftovers only after checking that size, timestamp and hash match (hashing over S: is slow, so this suits a handful of leftovers, not a whole tree):
$src = '<source>'; $dst = '<dest>'
Get-ChildItem $src -Recurse -File -Force | ForEach-Object {
  $e = Get-Item ($dst + $_.FullName.Substring($src.Length)) -Force -ErrorAction SilentlyContinue
  if ($e -and $e.Length -eq $_.Length -and $e.LastWriteTime -eq $_.LastWriteTime -and
      (Get-FileHash $_.FullName).Hash -eq (Get-FileHash $e.FullName).Hash) { Remove-Item $_.FullName -Force -Verbose }
  else { "KEEP (no matching copy): $($_.FullName)" }
}
  1. /MOV leaves the empty folder tree behind. Remove it only if it holds no files:
if (-not (Get-ChildItem <source> -Recurse -File -Force)) { Remove-Item <source> -Recurse -Force }

Adding a meal vendor or plan

When a vendor or plan is added to moveFTP.ps1:

  1. Give its S:\...\referrals → C:\FTP\...\referrals copy-back line /maxage:90, like the existing ones:
robocopy S:\HealthAlign\prod\ftp\<plan>\<vendor>\referrals C:\FTP\<plan>\<vendor>\referrals /e /j /XD incoming /maxage:90
  1. A new vendor name must also be added to $vendors in C:\FTP\archiveMeals.ps1. New plans are picked up automatically (C:\FTP\*).
  2. Edit moveFTP.ps1 between task runs (the task starts on the 10s), keep a dated backup next to it (moveFTP - <date>.ps1, the existing convention), and keep it UTF-8 without a BOM. Check the next run’s result is 0.

Where archived data lives

Path Contents
E:\archive\iis-logs-pre-2025-09-24 IIS logs from before the server was restored (~40 GB)
E:\archive\auth-logs-pre-2025-09-24 auth.myhomealign.com app logs (~90 GB)
S:\Archive\haweb01-c-documents-pre-2025-09-24 Old HealthAlign Documents\Provider and Visit copies (~221 GB). The live app reads C:\data\webapp\<tenant>\Documents inside its containers, which is S:\HealthAlign\<env>, not these.
S:\Archive\meals\<plan>\<vendor>\referrals Meal referral files older than 90 days (grows daily)
E:\archive\*.log robocopy logs for each move, and meals-archive.log for the daily archive

S: has no file-level backup or snapshots yet, so anything deleted there is gone.

Monitoring

The alert low disk space - thb-p-sftp-win (infra/ha-infra/business_unit_1/production/sftp_disk_alert.tf) checks every 15 minutes and fires one alert per drive when C: or E: averages below 10% free. It goes to PagerDuty Azure Alerting (High) and from there to #triage. Data comes from the Azure Monitor Agent into THB-PROD-LAW-VM-Insights (InsightsMetrics, Computer = HAWEB01).

To test it end to end, edit the rule’s condition in the portal from FreePercent < 10 to < 90, wait for the PagerDuty incidents (up to 15 minutes), then change it back to exactly < 10. Until it’s reverted, Terraform shows drift on the rule, and the next ha-infra apply resets it.

Gotchas

  • A Cerberus log can show 0 bytes in a folder listing while it’s open. Read the file (Get-Content -Tail) before assuming Cerberus stopped logging.
  • Don’t run cleanmgr from a scheduled task. Without a saved /sageset profile, cleanmgr /sagerun hangs in session 0, and the task stays “Running” and skips its next runs.
  • After freeing space, read free space again rather than trusting folder sizes. Folder totals have overstated what a move frees before.
Edit this page