What it is¶
wget downloads files over HTTP(S) and FTP, follows links recursively, and resumes interrupted transfers. It suits batch and background downloading; for API work and header debugging, use curl.
Basics¶
wget https://example.com/file.iso
wget -c https://example.com/file.iso # continue a partial download
wget -O out.iso https://example.com/file.iso # choose the output name
wget -b -o wget.log https://example.com/file.iso # run in background, log to file
Credentials¶
wget --user=username --ask-password https://host.example.com/private/file.txt
--ask-password prompts instead of putting the password on the command line (where it shows up in ps and shell history). --password='...' works but is not recommended.
Downloading images (or any type) from a page¶
wget -r --no-parent -A jpg https://host.example.com/photos/
A more polite version that skips files matching a pattern, resumes, throttles, and logs:
wget -r --no-parent -A jpg -R '*a.jpg' -c --limit-rate=20k -o "$HOME/wget.log" \
https://host.example.com/photos/
| Option | Meaning |
|---|---|
-r |
recursive |
--no-parent |
never ascend above the starting directory |
-A jpg |
accept only these suffixes (comma-separated list allowed) |
-R '*a.jpg' |
reject files matching the pattern |
-c |
resume partial files |
--limit-rate=20k |
cap bandwidth at 20 KB/s |
-o file |
write log output to file |
-w 1 --random-wait |
pause between requests to be kind to the server |
Mirroring a site¶
wget --mirror --convert-links --page-requisites --no-parent https://example.com/docs/
Respect robots.txt and the site's terms; hammering a server can get you blocked.