Build a Labeled Image Dataset From Google Images Search Terms
Turn one search term per class into a labeled image manifest with Google Images Scraper: full size URLs, width, height and the source page to check.
The actor referenced in this article. Pay only for results delivered.
Training or testing an image classifier starts with a few hundred labeled pictures per class. Google Images already sorts the web’s pictures by what they show, so one search term per class is a fast way to collect candidates. By hand, that means saving files one at a time, losing track of where each came from, and having no record of the page you need to check the license on.
Try it live: Google Images Scraper: Full Size Image URLs, Sizes, Sources. Pay per result delivered. Failed and empty results are never charged.
The Google Images Scraper returns one row per image for each search term, with the full size file URL, its width and height, and the page it was found on. The search term becomes your label, and the source page becomes your license trail.
What you get and who it is for
| Field | What it does in a dataset |
|---|---|
query | The search term, which you map to a class label |
image_id | Google’s id for the image, the same in every run, used to deduplicate |
image_url | The full size file on the publisher’s site |
image_width, image_height | Size in pixels, to drop small images before download |
thumbnail_url | Google’s small copy, useful for a quick visual review |
source_page_url | The page to open and check the license on |
source_domain, source_name | The site, for spotting sources you want to exclude |
dominant_color | Main color as hex, handy for spotting off-class images in bulk |
This is for machine learning engineers bootstrapping a classifier, researchers who need a small evaluation set, and teams building reference sets for visual QA. It is a candidate collector, not a finished dataset: you still review labels and licenses. The work in between, dedupe, review and storage, is what a scraping pipeline does between raw capture and a clean dataset.
The input
{
"queries": [
"golden retriever puppy",
"golden retriever puppy playing",
"beagle puppy",
"beagle puppy playing",
"dachshund puppy",
"dachshund puppy playing"
],
"maxResultsPerQuery": 200,
"country": "us",
"language": "en",
"imageSize": "large",
"imageType": "photo",
"usageRights": "creativeCommons",
"safeSearch": true
}
queries: one or more search terms per class. Two phrasings per breed widen the pool, and the script maps both to one label.imageSize: largeandimageType: photoare Google’s own Size and Type filters. They cut clip art and small files, but Large is Google’s definition, so the script also enforces a minimum side length.usageRights: creativeCommonskeeps images Google believes carry a Creative Commons license. It is a filter, not a license. More on that below.safeSearch: truefilters explicit images, which you want in any training set.
What the output looks like
[
{
"query": "golden retriever puppy",
"position": 1,
"image_id": "jUS0xT_ulT-SUM",
"image_url": "https://s3.amazonaws.com/cdn-origin-etr.akc.org/wp-content/uploads/2020/07/09151754/Golden-Retriever-puppy-standing-outdoors.jpg",
"image_width": 729,
"image_height": 486,
"source_page_url": "https://www.akc.org/expert-advice/dog-breeds/golden-retriever-puppy-training-timeline/",
"source_domain": "akc.org",
"source_name": "American Kennel Club",
"file_size": "190KB",
"dominant_color": "#b0a06a"
},
{
"query": "golden retriever puppy",
"position": 2,
"image_id": "aZUu7S9Ewduo5M",
"image_url": "https://goldenmeadowsretrievers.com/wp-content/uploads/2023/08/Golden-Retriever-Puppies-Are-Simply-Irresistible-scaled.webp",
"image_width": 2560,
"image_height": 1271,
"source_page_url": "https://goldenmeadowsretrievers.com/golden-retriever-puppies-are-simply-irresistible/",
"source_domain": "goldenmeadowsretrievers.com",
"source_name": "Golden Meadows Retrievers",
"file_size": "202KB",
"dominant_color": "#f8ebcb"
}
]
Both rows come from our run on golden retriever puppy on October 5, 2026, with no filters set, so they show the unfiltered grid. Of the 20 rows in that run, 18 had a shorter side of at least 512 pixels and 14 were at least 1,000 pixels wide. Six of the 20 came from Instagram, Facebook, Pinterest and Reddit, which is the kind of source the usage rights filter is meant to cut. When you set filters, each row also carries a filters field recording them.
Run it in Apify Console
- Open https://apify.com/themineworks/google-images-scraper and click Try for free.
- In Search terms, enter your terms, one per line.
- Set Images per search term to 200, Size to Large and Type to Photo.
- In Usage rights, pick the Creative Commons option. Turn on SafeSearch.
- Click Start and watch the rows in the Output tab.
- Click Export and download CSV, JSON or Excel. The CSV is a ready manifest: sort by
query, then review thumbnails and source pages.
Run it from Python
pip install apify-client
import csv
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
# search term -> class label; several phrasings can share one label
LABELS = {
"golden retriever puppy": "golden_retriever",
"golden retriever puppy playing": "golden_retriever",
"beagle puppy": "beagle",
"beagle puppy playing": "beagle",
"dachshund puppy": "dachshund",
"dachshund puppy playing": "dachshund",
}
MIN_SIDE = 512
FIELDS = ["label", "image_id", "image_url", "width", "height",
"source_page_url", "source_domain", "license", "license_checked"]
run = client.actor("themineworks/google-images-scraper").call(run_input={
"queries": list(LABELS),
"maxResultsPerQuery": 200,
"imageSize": "large",
"imageType": "photo",
"usageRights": "creativeCommons",
"safeSearch": True,
})
seen = {} # image_id -> label
ambiguous = set()
rows = []
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
if item.get("_type") == "info":
continue # the run summary row, never billed
label = LABELS.get(item["query"])
if label is None:
continue
if min(item["image_width"], item["image_height"]) < MIN_SIDE:
continue
image_id = item["image_id"]
if image_id in seen:
if seen[image_id] != label:
ambiguous.add(image_id) # same picture under two labels
continue
seen[image_id] = label
rows.append({
"label": label,
"image_id": image_id,
"image_url": item["image_url"],
"width": item["image_width"],
"height": item["image_height"],
"source_page_url": item["source_page_url"],
"source_domain": item["source_domain"],
"license": "", # fill in after checking the page
"license_checked": "", # write yes once confirmed
})
rows = [r for r in rows if r["image_id"] not in ambiguous]
with open("manifest.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=FIELDS)
writer.writeheader()
writer.writerows(rows)
per_label = {}
for r in rows:
per_label[r["label"]] = per_label.get(r["label"], 0) + 1
print(per_label)
print(f"{len(ambiguous)} images dropped for appearing under two labels")
The script maps each term to its label, drops images smaller than 512 pixels on the short side, deduplicates by image_id, removes any image that turned up under two labels, and writes manifest.csv with two empty columns for your license check. Once you have opened the source pages and marked rows yes, this downloads only the cleared files into one folder per label:
import csv
import os
import urllib.request
from urllib.parse import urlparse
with open("manifest.csv", encoding="utf-8") as f:
cleared = [r for r in csv.DictReader(f) if r["license_checked"] == "yes"]
for r in cleared:
ext = os.path.splitext(urlparse(r["image_url"]).path)[1] or ".jpg"
folder = os.path.join("dataset", r["label"])
os.makedirs(folder, exist_ok=True)
request = urllib.request.Request(r["image_url"], headers={"User-Agent": "dataset-builder"})
try:
with urllib.request.urlopen(request, timeout=20) as resp:
data = resp.read()
with open(os.path.join(folder, r["image_id"] + ext), "wb") as out:
out.write(data)
except Exception as err:
print("skipped", r["image_id"], err)
Refresh it every month
Datasets go stale as new pictures appear. Save the input as a task in Apify Console, add "timeRange": "month" so Google returns only images it found in the past month, then open Schedules and create a schedule with the cron expression 0 7 1 * * (07:00 on the first of each month) for that task. Because image_id is stable across runs, append only ids that are not already in your manifest. Apify can send each finished run to a Google Sheet or call a webhook, so new candidates can land in a review queue instead of your inbox. The actor has no separate “only new” switch; the time filter plus the id check does that job.
What it costs
You pay per image delivered: $2.09 per 1,000 on the Bronze plan, $1.79 on Silver, and $1.49 on Gold, Platinum and Diamond, plus a flat $0.005 per run start.
The input above asks for up to 1,200 images (six terms at 200), which is at most $2.51 on Bronze plus $0.005. The Creative Commons filter usually leaves Google with fewer images than that, and you pay only for what is delivered, so treat $2.51 as the ceiling. Duplicates within a term are never charged.
Limits worth knowing
- The usage rights filter is not a license. Google reads license details from the pages it indexes, and it can be wrong or out of date. Creative Commons licenses also differ: some require attribution, some forbid commercial use or changes. Open
source_page_urlfor every image you keep and record the license. - Labels come from the search, not from looking at the picture. Google returns what it ranks for the words, so expect off-class images in every class and plan a review pass using
thumbnail_url. image_idcatches exact repeats only. In our golden retriever puppy run, positions 8 and 12 were two different ids from the same Instagram post, at the same size and file size. Near-duplicates need a perceptual hash after download.- Some publishers refuse downloads from other sites. Expect a share of
image_urlfetches to fail. The thumbnail is Google’s copy and much smaller. - Google stops at a few hundred per term. In the actor’s own test Google served 355 images for red sneakers and then an empty page. Filters cut that further, so add phrasings rather than raising
maxResultsPerQuery.
Related
Explore the scraper referenced in this article: inputs, outputs, and pricing, then run it on Apify.
Frequently asked questions
Does the Creative Commons filter mean I may use the images? +
No. It narrows results to images Google believes carry a Creative Commons license, based on what Google reads from the page. The image still belongs to its owner, so open source_page_url and confirm the license and its terms before you use the file.
Does the actor download the image files? +
No. It returns links, sizes and source pages, not the files. You download the files you have cleared yourself, for example with the second script in this guide.
How do I get more images for one class? +
Add several phrasings that mean the same class, such as golden retriever puppy and golden retriever puppy playing, and map them to one label. Each term is searched on its own, and deduplicating by image_id removes the overlap.
Can the same image show up under two labels? +
Yes, because each search term is searched separately. image_id is Google's id for the image and stays the same across terms and runs, so drop any id that appears under two different labels.
Track Which Sites Rank in Google Images for Your Keywords
Use Google Images Scraper to see which domains hold the top image positions for your keywords in each country, then rerun weekly to catch movement.
Pull the Top 100 Best Sellers for Every Amazon Subcategory in One Run
Walk an Amazon department with Amazon Bestsellers Scraper and get the top 100 of every subcategory in one run, with a hard cap on lists and cost.
Benchmark LinkedIn Post Engagement Across Company Pages
Compare reactions and comments per post across LinkedIn company pages, scaled by follower count, with the LinkedIn Company Posts Scraper. No login.
Build a List of Funded Startups That Are Hiring on Wellfound
Turn Wellfound job listings into one row per startup with website, stage, total raised and latest round, using the Wellfound Jobs Scraper.