A web server administrator notices that `top` shows high 'wa' (I/O wait) percentage, suggesting a disk bottleneck. However, `iostat -x` reports low `%util` (e.g., 20%) for all disks. The server is running a database that frequently accesses small files. What is the most likely reason for this discrepancy?
- AThe server is experiencing network congestion, which `iostat` does not measure.
- BThe `iostat` command is not running with sufficient interval to capture bursty I/O.
- CThe `top` command is misinterpreting CPU wait time, and the issue is actually CPU contention.
- DThe disk bottleneck is due to a high number of small I/O operations (IOPS), not high data throughput.
Show answer & explanationAnswer & explanation
Correct answer: D. The disk bottleneck is due to a high number of small I/O operations (IOPS), not high data throughput.
This scenario describes a classic 'IOPS bottleneck' where the disk is busy with many small, rapid I/O operations rather than large, sustained data transfers. `%util` measures the percentage of time the device is busy, which can be low if the *volume* of data is small, even if the *number* of operations (IOPS) is very high, causing latency. `top`'s `wa` accurately reflects the CPU waiting for these numerous, albeit small, I/O operations to complete. `iostat`'s `r/s` and `w/s` (reads/writes per second) or `await` (average I/O wait time) would likely be high in this case, even if `rkB/s` and `wkB/s` (kilobytes per second) are low, leading to low `%util` but high `wa`.
Why the other options are wrong
- A. Network congestion would lead to high network latency, but `top`'s 'wa' specifically indicates CPU waiting for *disk* I/O, not network I/O. `iostat` *does* measure disk activity.
- B. While `iostat` interval matters for bursty I/O, it wouldn't explain consistently low `%util` if the disk is truly busy with high IOPS. The fundamental issue is how `%util` is derived.
- C. `top`'s 'wa' is generally accurate for CPU I/O wait. If the issue were CPU contention, `%us` or `%sy` would be high, not `%wa`.
IOPS vs. Throughput Bottleneck
A disk can be bottlenecked by either low data throughput (low MB/s) or low I/O operations per second (IOPS). `%util` in `iostat` might be low if throughput is low, but high `r/s` or `w/s` with high `await` can still cause high CPU I/O wait ('wa' in `top`) due to many small, latent operations.
- IOPS (I/O Operations Per Second) measures frequency of I/O.
- Throughput measures data volume (MB/s).
- High IOPS with small files can cause high 'wa' even with low `%util` if `await` is high.
Memory trick: IOPS can hide, while UTIL lies, but WAITING still makes the CPU cry.