GNU ddrescue said 100.00%. My first structured file-recovery pass produced 0 useful recovered user files.
Those two results came from the same 4 TB hard-drive failure, and the contradiction is the reason this recovery became much more interesting than “clone the bad disk and copy the files.” ddrescue had done its job remarkably well: it had copied approximately 99.999654% of the physical source. The clone was on healthy hardware. Yet APFS was still corrupt, macOS would not give me the filesystem normally, and the first APFS recovery stack could enumerate large parts of the directory tree while failing to read the contents of ordinary files.
I eventually recovered every selected, enumerated folder tree I needed, but only after treating the incident as three separate problems: physical block recovery, damaged-filesystem interpretation, and large-scale file extraction with validation and resume support.
The failure stopped being a filesystem problem first
The original 4 TB Toshiba had reached the point where I no longer trusted it with normal filesystem activity. Some reads took 60–75 seconds. Operations could hang. The drive intermittently disappeared from macOS, clicked audibly, and sometimes powered down.
Shortly before the failure, I had written roughly 300 GB of additional data and performed a mass rename affecting around 500,000 files and directories. The timing made that metadata-heavy workload suspicious, but I cannot prove that it caused the hardware failure. It may simply have stressed an already unhealthy drive enough to expose the problem.
What I could establish was the drive’s behavior. Once mechanical storage is stalling, disappearing, and clicking, repeatedly browsing directories is the wrong abstraction. Directory traversal can trigger more reads and seeks. Mounting a filesystem can trigger metadata work. Every experiment consumes time on the only component whose remaining useful life is unknown.
So I changed the objective from:
recover my files
to:
recover as many readable sectors as possible
I used GNU ddrescue 1.30 with a persistent mapfile. The exact physical size of the source drive was:
4,000,787,027,968 bytes
The mapfile was essential because the source was not stable enough for a one-shot copy. It let the rescue survive stalls, disconnects, restarts, and later passes without forgetting which regions had already been recovered.
I also hit a practical issue with the raw macOS device path: the apparent rescue extent could become nonsensical instead of ending at the real device boundary. I therefore constrained the rescue domain to the known physical size above. In this case, explicitly bounding the input was a correctness measure, not a speed optimization.
Why ddrescue showing 100.00% did not mean the files were safe
Near the end of the physical rescue, ddrescue reported approximately:
domain size: 4000 GB
rescued: 4000 GB
non-tried: 13481 kB
non-trimmed: 327680 B
non-scraped: 0 B
bad-sector: 27136 B
The headline percentage was:
100.00%
But the unresolved states still totaled:
13,835,816 bytes
or roughly:
13.84 MB
Against a 4,000,787,027,968-byte source, the recovered share was approximately:
99.999654%
That is an excellent block-recovery result. It is not a file-integrity result.
The location of missing bytes matters more than the headline quantity. Several megabytes lost from unused space may affect nothing visible. A small unreadable region inside a video may damage one file. A much smaller loss in filesystem metadata can make many otherwise intact data extents difficult to locate.
This became the central mental model for the rest of the recovery:
| Layer | Question it answers | What success does not prove |
|---|---|---|
| Block recovery | Were physical sectors copied? | That APFS can reconstruct every file |
| Filesystem recovery | Can paths, metadata, and extents be resolved? | That every extracted byte is valid |
| File validation | Did a file arrive at the expected size or hash? | That no undiscoverable file ever existed |
I made a few final attempts against the remaining unread regions. They eventually stopped producing useful new reads while the Toshiba was clicking heavily. That was where I stopped using the original as an active recovery source.
Why I bought two 5 TB drives to recover one 4 TB disk
The first new drive was a 5 TB Seagate Expansion with an exact physical capacity of:
5,000,981,077,504 bytes
I wrote the block-level Toshiba clone to it. The source layout occupied roughly the first 4 TB, leaving about 1 TB beyond the copied layout. I deliberately left that extra capacity alone.
I did not enlarge the APFS container. I did not repartition the clone for convenience. I did not run a filesystem repair on it. That disk became the master clone.
Then I bought a second 5 TB drive. It was freshly formatted, writable, independently tested, and used only for recovered output.
failing 4 TB HDD
│
│ GNU ddrescue
▼
5 TB drive #1
master block-level clone
READ-ONLY
│
│ APFS parsing and extraction
▼
5 TB drive #2
recovered files
WRITABLE
This meant buying roughly 10 TB of nominal new storage to recover a volume containing about 3.26 TB of used data. The extra disk was not about capacity. It was about preserving an invariant:
If an experiment is wrong, I can return to the same untouched master clone.
Repairing, resizing, repartitioning, or writing recovered output onto the master would have mixed preservation with experimentation. Keeping source and destination on separate physical drives made mistakes recoverable.
Before trusting the destination, I ran an approximately 10 GB write/read test. In that test it sustained roughly 144.4 MB/s in both directions. Sample raw reads from the master clone were roughly 28–49 MB/s and did not reproduce the original drive’s physical I/O failure pattern during those checks.
At that point, the problem had changed. I was no longer debugging failing hardware. I was debugging damaged APFS metadata preserved on hardware that behaved normally in my tests.
The APFS clone was readable as a device but invalid as a filesystem
The cloned APFS physical store was:
4,000,650,887,168 bytes
The relevant partition began at sector:
264192
With 512-byte sectors, that is a byte offset of:
135,266,304 bytes
The APFS volume reported approximately:
3,260,976,717,824 bytes
consumed.
A read-only APFS check eventually reached the fsroot tree and reported:
Checking fsroot tree.
error: (oid 0xe36cb) apfs_root:
btn: dev_read_finish(3828538, 1): Input/output error
fsroot tree is invalid.
This was the point where “ddrescue copied almost the whole disk” stopped being useful as a complete diagnosis. The raw clone existed. The filesystem structure inside it was still inconsistent.
I deliberately kept filesystem checking non-modifying. I did not convert fsck_apfs -n into a repair operation against the only high-quality master clone I had. A repair can be appropriate on ordinary storage, but here it would have changed the evidence I was still trying to understand.
The Sleuth Kit could list the APFS namespace but failed on file contents
The Sleuth Kit 4.15.0 was the first recovery stack that made the clone look promising. Using the whole cloned device, the known partition offset, and the APFS superblock I had identified, I could enumerate real directory names with fls:
fls \
-o 264192 \
-B 4594668 \
-p \
/dev/rdiskN
I use N deliberately. macOS disk numbers changed across reconnects and reboots, so I did not treat a previous /dev/disk6 assignment as identity.
fls could traverse substantial parts of the namespace. One large tree alone exposed roughly 8,900 directories. For a moment, that looked like the hard part was solved.
Then I tried to retrieve file contents.
A representative failure was:
libc++abi: terminating due to uncaught exception
of type std::runtime_error:
could not read APFSBlock
tsk_recover could begin extracting and then fail on an APFS block. Individual icat attempts showed the same class of problem.
The important distinction was:
directory traversal works
did not imply:
file content retrieval works
A parser can have enough surviving metadata to discover a pathname while still failing later when resolving the file object, extent metadata, or content blocks needed to return the byte stream.
I made the first bulk recovery engine fault-tolerant, but the parser was still wrong for this damage
My first response was to make TSK extraction more resilient instead of immediately changing parsers.
I built a Python recovery wrapper around fls and icat. It kept durable state in SQLite, logged failures, supported resume, wrote partial output separately, and deprioritized low-value macOS housekeeping data during the first pass.
I also added a “hotspot” rule: if four consecutive files in one directory failed with the same zero-byte APFSBlock pattern, the script stopped spending time on the rest of that branch, deferred it, and continued elsewhere. The working hypothesis was that a cluster of identical failures might share one damaged metadata dependency rather than represent hundreds of independently destroyed payloads.
The orchestration was useful. The underlying APFS reader was not.
At one captured point, the recovery database contained:
DEFERRED_HOTSPOT: 30,356 files
FAILED: 486 files
The 486 failures broke down into:
409 APFSBlock crashes
75 rc=0 but output-size mismatch
2 other rc=1 failures
And the structured pass had produced:
0 useful recovered user files
The 75 size mismatches exposed a bug in my own wrapper: some system metadata files had returned data, but my parser had recorded an expected size of zero. Fixing that interpretation mattered, but it did not change the dominant result. Ordinary user files were still ending in zero bytes with could not read APFSBlock.
I picked one small AVIF file as a reproducible test case. TSK knew its expected size:
expected size: 56,309 bytes
recovered: 0 bytes
That file became much more valuable than another multi-hour bulk pass. If a new approach could not recover one 56 KB test file that failed consistently, it did not deserve access to the remaining terabytes.
The disk blocks and the APFS metadata path were different failure domains
By this point I had three observations:
GNU ddrescue:
almost the entire physical source was copied
TSK fls:
many real paths were discoverable
TSK icat:
many ordinary file contents still failed
Those observations are compatible once the lookup chain is separated conceptually:
pathname
↓
directory record
↓
file object / inode metadata
↓
extent mapping
↓
physical data blocks
Data blocks near the bottom can survive while a link higher in the metadata chain is damaged. A different possibility is that two APFS implementations traverse the same damaged structures differently.
I did not isolate one corrupt APFS object that explained every failure, so I would not claim that as a confirmed root cause. What the evidence did support was a much more useful experiment: keep the cloned bytes unchanged and change the parser.
142 APFS checkpoints did not become an instant rollback path
A raw, read-only scan of the APFS checkpoint descriptor area found 142 candidate NXSB checkpoint superblocks, with transaction IDs ranging from:
223133
down through:
222992
An obvious question was whether an older checkpoint referenced a healthier metadata tree.
I built apfs-fuse and tried different checkpoint transaction IDs through fuse-t. The newest checkpoint stalled. Older XIDs did the same. I added hard per-checkpoint timeouts, and more than 50 consecutive attempts failed to produce a usable mount. I also tried both NFS and SMB fuse-t backends.
That did not prove all of the checkpoints were corrupt. The experiment depended on the checkpoint, apfs-fuse, fuse-t, macOS device behavior, and the mount backend. A failure anywhere in that stack could produce the same visible outcome.
A later parser reported zero APFS snapshots, which also reinforced a distinction I had to keep straight: those checkpoint transaction states were not the same thing as user-visible APFS snapshots.
A parser that hangs on --help cannot diagnose my disk
I also built go-apfs-v2. The build completed, but even:
apfs --help
hung and had to be killed by a timeout.
Minimal block inspection against both raw and buffered device paths hung as well. I made the same kind of bounded test with apfsutil; both device forms timed out after about 15 seconds.
Those were useful failures because they stopped me from drawing the wrong conclusion. A tool that cannot reliably complete its own help path is weak evidence about whether a damaged file is recoverable.
My rule became:
Validate the recovery tool before treating the recovery tool’s failure as evidence about the data.
Preventing macOS from auto-mounting the clone removed another source of risk
macOS itself was another moving part. I wanted the healthy destination mounted normally, but I did not want Disk Arbitration automatically trying to mount the damaged APFS clone whenever I reconnected storage.
The safe sequence I ended up using was:
- Connect and mount the healthy recovery destination.
- Verify its volume identity and free space.
- Freeze
diskarbitrationd. - Verify there is no active
mount_apfsprocess. - Attach the master clone.
- Identify the master dynamically by known physical size and APFS identity.
- Verify the master is not mounted.
- Perform read-only recovery work.
- When stopping, save state, resume Disk Arbitration, and then shut down or eject normally.
The command I actually used to freeze Disk Arbitration was:
sudo kill -STOP "$(pgrep -x diskarbitrationd)"
I verified the process state contained T, and I checked for stray mount_apfs processes before continuing.
The important safety lesson was not the specific disk number. It was the opposite: never trust yesterday’s disk number. After a reboot, a previous /dev/disk6 can refer to a different physical device. I treated the APFS UUID and known physical size as identity, then derived the current device path from that.
libfsapfs recovered the same file TSK returned as zero bytes
The breakthrough came from libfsapfs, a different APFS implementation.
I built it from source on macOS. The resulting fsapfsinfo binary identified itself as:
fsapfsinfo 20260923
Then an important macOS device-interface difference appeared.
The raw character-device partition:
/dev/rdiskNs2
failed quickly with an invalid-argument read near offset 4096.
The buffered block-device form:
/dev/diskNs2
worked.
Same physical clone. Same APFS partition. Different macOS device interface.
fsapfsinfo opened the container and found one volume. I then tested the exact 56,309-byte file that TSK had failed to extract.
TSK had produced:
expected: 56,309
recovered: 0
APFSBlock failure
Through libfsapfs, the file entry produced:
size: 56,309
MD5: c6f56db33eafc1de0f52a035bc255dc7
RC: 0
This was the first result that materially changed the diagnosis. The cloned bytes had not changed. The test file had not changed. The damaged APFS had not been repaired.
The APFS implementation had changed.
At minimum, I now knew that a real user file which looked unrecoverable through TSK was still reachable deeply enough through libfsapfs to read its full contents and calculate a digest.
I extracted one known-bad file before trusting libfsapfs with terabytes
A successful digest was not enough for me to launch a multi-terabyte recovery. I wanted the actual bytes on the destination disk.
Instead of guessing the libfsapfs C API, I inspected the library source and followed the read path already used by fsapfsinfo when calculating the digest. I then built a small read-only extractor for the one known test file.
The extractor wrote exactly:
56,309 bytes
to the separate recovery disk. The extracted file’s MD5 was:
c6f56db33eafc1de0f52a035bc255dc7
That matched the earlier digest.
Only then did I scale the approach. The one-file test had proven three separate things: the pathname could be resolved, the complete expected byte count could be extracted, and the extracted bytes produced the same digest as the earlier full-content read.
The bulk recovery problem was mostly about fault containment and resume
Once libfsapfs could recover a file TSK could not, the hard problem changed again. I needed a system that could process a very large directory tree without one damaged branch, one reboot, or one Ctrl+C turning the job into a restart from zero.
The bulk recovery pipeline therefore used a few strict invariants:
- The master clone was opened read-only.
- Recovered output was written only to the second 5 TB disk.
- Directory and file names were preserved.
- Progress lived in SQLite so it survived process exits and reboots.
- Each file was written to a temporary path first.
- A temporary file was renamed to its final path only after the full expected size had been written.
- Existing files of the expected size could be recognized during resume.
- Failed or problematic work was kept separate from completed work.
- Free space was checked and a safety reserve was maintained.
- Regeneratable development artifacts and low-value system metadata could be skipped or deprioritized.
One performance change mattered immediately: I stopped reopening the APFS container independently for every file.
The fast path processed one directory at a time. A worker opened the source, enumerated that directory, recovered its immediate regular files, and returned child directories to the queue. If a directory failed or timed out, the controller marked it DEFERRED and moved on instead of blocking the global pass.
After no normal pending work remained, deferred directories were revisited through a slower fallback path with more isolated per-file work. Remaining failures could then be retried independently.
discover directory
↓
recover immediate files
↓
verify expected sizes
↓
commit durable state
↓
queue child directories
↓
defer local failures
↓
continue globally
↓
fallback and retry later
This architecture matched the actual failure pattern much better than one giant recursive command. Damage was not uniform, so the recovery system should not make progress uniform either.
Why I used expected-size checks and atomic rename instead of hashing millions of files
MD5 was useful during the single-file proof because I needed strong evidence that the parser was reading the complete content of a file that TSK could not.
Doing a second full read of every recovered byte just to hash millions of files would have added a large amount of I/O. For the main extraction pass, I used a different invariant.
For each regular file, APFS metadata supplied an expected size. The worker wrote to a temporary file and only promoted it to the final pathname after the complete read matched that expected size.
That means an interruption should not leave a short file masquerading under the final name.
A size match is not a cryptographic integrity proof. I do not treat it as one. But expected-size verification plus atomic rename was a practical correctness boundary for the high-volume pass, while targeted hashes remained useful for samples and known failure cases.
SQLite made a reboot boring instead of catastrophic
The recovery ran long enough that I needed to stop the computer and resume later. That requirement changed the design from “script” to “recoverable workflow.”
A Ctrl+C did not simply abandon the active child process. The controller trapped the interruption, stopped the worker, returned the active directory to a recoverable state, committed SQLite, and exited.
A clean stop looked conceptually like:
CURRENT DIRECTORY -> PENDING
CTRL+C: RECOVERY STOPPED SAFELY
STATE SAVED. RUN THE SAME COMMAND TO RESUME.
After rebooting, I repeated the disk-identity checks and started the same recovery command. The state database resumed the existing queue.
One later restart began with:
DEFERRED: 1
DONE: 73,015
PENDING: 7,254
That was much more meaningful than a generic progress bar. It showed that tens of thousands of completed directory units had survived the reboot and that the remaining queue was explicit.
On an earlier resume, the directory interrupted in the previous session was picked up again. Files already present at the expected size were recognized, and only the missing work had to be written. That was the behavior I wanted: restarting recovery should be routine, not frightening.
326,799 files were the first proof that the method scaled
Before expanding the new extractor across all selected data, I used one large priority tree as a bulk validation target.
The completed state reported:
directories completed: 8,940
new files written: 302,541
existing/resumed files: 24,258
recorded failed files: 0
The destination contained:
326,799 files
145,039,215,948 bytes
or about:
135.08 GiB
The file counts reconciled exactly:
302,541 + 24,258 = 326,799
There were no leftover .partial.* files in that completed tree.
The 56,309-byte test file had proved the parser could succeed where TSK failed. The 326,799-file recovery proved that the same approach could survive a substantial real hierarchy with resume logic, existing-file detection, and no recorded file failures in that completed pass.
The larger recovery crossed 1.6 million new files before completion
After the priority tree completed cleanly, I expanded the recovery to the remaining selected top-level data.
At one intentional safe stop, SQLite reported:
DONE directories: 39,015
PENDING directories: 17,017
DEFERRED directories: 1
new files written: 1,636,305
new bytes written: 902,715,716,335
That was approximately:
840.72 GiB
of newly written data recorded by the directory passes at that point.
The single DEFERRED directory did not mean lost data. It meant the fast path had intentionally stopped letting that local problem delay unrelated work. The fallback phase existed specifically to revisit such cases later.
Subsequent sessions resumed from the same database. The completed count increased and the pending queue decreased. In the end, I recovered every selected, enumerated folder tree I needed.
What I can honestly claim about the final recovery
I will not describe the outcome as “every byte recovered.” The evidence does not support that.
The original ddrescue map still contained roughly 13.84 MB that had not been confirmed as successfully copied. I also cannot prove that no filesystem object became completely undiscoverable because the metadata needed to enumerate it was among the damaged regions.
Those limitations matter because structured recovery can prove that an enumerated object was extracted; it cannot prove the historical non-existence of an object the damaged namespace can no longer reveal.
The strongest final statement is narrower:
Every selected, enumerated folder tree I needed was recovered successfully through the structured recovery process.
I did not need to repair the master clone in place. I did not need to reuse the original clicking Toshiba for the bulk extraction. I did not need a whole-disk PhotoRec-style carving pass that would sacrifice directory structure and filenames.
The failed tools were still useful evidence
In retrospect, the successful path sounds simple:
ddrescue clone
↓
libfsapfs
↓
resumable extractor
↓
recovered files
That is not how the investigation felt while it was happening, and removing the failed approaches would remove much of the useful engineering lesson.
TSK taught me that APFS namespace traversal and content retrieval were different failure modes.
My first wrapper bug taught me that a recovery script’s own error classification is not ground truth.
The hotspot mechanism taught me to isolate local failure instead of allowing it to stall global progress.
The checkpoint experiment taught me not to confuse APFS transaction checkpoints with snapshots.
The FUSE attempts taught me that a failed mount can implicate several layers besides the filesystem data itself.
The APFS reader that hung on --help taught me to validate the tool before interpreting its diagnostics.
The /dev/rdiskNs2 versus /dev/diskNs2 behavior taught me that the operating-system I/O path can change a parser’s behavior even when the underlying disk is identical.
And the two-drive architecture gave me the freedom to make mistakes everywhere else while leaving the master clone unchanged.
The recovery workflow I would use again
- Stop ordinary filesystem activity on mechanically failing storage. If reads are stalling, the device disappears, or it clicks, I would prioritize a resumable block-level clone over Finder exploration.
- Use GNU ddrescue with a persistent mapfile and a verified rescue domain. The mapfile preserves progress; a known source size prevents device-size confusion from becoming part of the recovery.
- Keep one master clone read-only. Do not repair it, resize it, repartition it, or use it as recovered-file storage.
- Write recovered files to a second physical disk. Source preservation and output storage are different jobs.
- Diagnose the cloned filesystem read-only first. A healthy physical clone containing corrupt APFS metadata is a logical recovery problem, not the same problem as a clicking source disk.
- Choose one reproducible failed file as a parser test. A 56 KB known-bad file told me more about alternative parsers than hours of blind bulk extraction.
- Validate the parser itself. If a tool hangs before it meaningfully reads the source, do not interpret that as proof that the data is gone.
- Do not assume one APFS implementation defines recoverability. TSK and
libfsapfsbehaved very differently on the same cloned bytes. - Identify disks by stable properties, not temporary device numbers. A filesystem UUID and known physical size are safer than yesterday’s
/dev/diskN. - Make long-running extraction resumable from the beginning. Durable state, temporary files, atomic rename, deferred work, bounded retries, and clean shutdown handling are part of correctness at this scale.
- Separate validation levels. A ddrescue percentage, a visible pathname, an expected-size match, and a content hash prove different things.
The number that looked like the finish line was only the end of stage one
The most misleading number in the whole recovery was still:
100.00%
It looked like an answer to “Did I save the disk?”
What it actually answered was much narrower:
How much of the physical rescue domain did ddrescue successfully copy?
It did not answer whether APFS could reconstruct the namespace. It did not answer whether one parser could traverse damaged metadata that another parser rejected. It did not answer whether a recovered file had the expected length. And it did not answer whether a multi-million-file extraction could survive errors and reboots without corrupting its own state.
I needed separate evidence for each layer.
The physical recovery succeeded first. The filesystem was still damaged. The first parser could expose names but failed on many file contents. A different APFS implementation successfully read the same test file. That proof became a one-file extractor, the extractor became a resumable recovery engine, and the engine eventually recovered the selected folder trees I needed onto a second disk while the master clone remained untouched.
The original hard drive never became healthy. APFS never magically repaired itself. What changed was the recovery model.
Block recovery, filesystem recovery, and file validation are separate engineering stages. Treating them as one problem made the situation look almost hopeless. Separating them made it tractable.