API WORKFLOWS

Web Content Extraction APIs

Turn an accessible public article into structured content for indexing, review, or a downstream AI workflow. Choose Readability for URL or raw HTML input, or Fetch Content for article text and page metadata from a URL.

A practical integration workflow

  1. Choose a public article you are permitted to access and process.
  2. Select URL or HTML input according to the endpoint contract.
  3. Inspect the extracted text and available metadata before storing or using it.

Choose an API and try an example

Each example uses the normal API endpoint. Replace YOUR_APPKEY with your own key. Use the separate Demo link to view fixed sample data without a subscription.

Readable Article Extraction

Submit a public article URL or HTML and extract readable HTML, text, and available metadata.

Input and coverage: Supply exactly one of url or html. URL maximum 2048 characters; HTML maximum 1 MiB in UTF-8. Some pages have no extractable article.

Example request

Replace YOUR_APPKEY with your own API key. This request uses the normal API endpoint.

curl --fail-with-body 'https://api.gugudata.io/v1/websitetools/readability?appkey=YOUR_APPKEY' \
  --request POST \
  --header 'Content-Type: application/json' \
  --data '{"url": "https://gugudata.github.io/gugudata-io/guides/article-content-extraction-api-seo-guide/"}'
View sample Demo response

This is a fixed Demo sample, shortened for readability. It is not the response to your example request. Your results depend on the input and dataset. File URLs are temporary and have been removed.

{
  "dataStatus": {
    "statusCode": 200,
    "status": "SUCCESS",
    "statusDescription": "successfully",
    "dataTotalCount": 1
  },
  "data": {
    "title": "Article Content Extraction API Integration Guide",
    "byline": "GuGuData.io",
    "dir": null,
    "lang": "en",
    "content": "<div id=\"readability-page-1\" class=\"page\"><div><h2 id=\"article-content-extraction-api-integration-guide\">Article Content Extraction API Integration Guide</h2>\n<p>Search teams, content platforms, and data products often start with the same messy input: a public article URL. The page may include navigation, cookie banners, related posts, comments, scripts, and layout HTML. What you usually need for an SEO workflow is much smaller and much more structured: the article title, description, readable body, plain text, main image, image candidates, author, published time, and source domain.</p>\n<p>The <a href=\"https://gugudata.io/details/fetchcontent\">GuGuData Article Content Extraction API</a> is designed for that exact job. It extracts clean article content from a public webpage URL and returns …",
    "textContent": "Article Content Extraction API Integration Guide\nSearch teams, content platforms, and data products often start with the same messy input: a public article URL. The page may include navigation, cookie banners, related posts, comments, scripts, and layout HTML. What you usually need for an SEO workflow is much smaller and much more structured: the article title, description, readable body, plain text, main image, image candidates, author, published time, and source domain.\nThe GuGuData Article Content Extraction API is designed for that exact job. It extracts clean article content from a public webpage URL and returns a normalized JSON response that can be stored, indexed, summarized, compared, or passed into downstream SEO automation.\nThis guide explains where the API fits in an SEO stack,…",
    "length": 10436,
    "excerpt": "Extract LLM-ready article text, readable HTML, metadata, images, author, publication time, and source for RAG, search, monitoring, and analysis.",
    "siteName": "GuGuData.io",
    "publishedTime": "2026-07-08T00:00:00.000Z"
  }
}

Webpage Content Extraction

Extract article text and metadata from a public webpage for search or content processing.

Input and coverage: Use a public HTTP(S) article URL. Page access restrictions and source structure can affect extraction; metadata may be absent.

Example request

Replace YOUR_APPKEY with your own API key. This request uses the normal API endpoint.

curl --fail-with-body 'https://api.gugudata.io/v1/websitetools/fetchcontent?appkey=YOUR_APPKEY' \
  --request POST \
  --header 'Content-Type: application/json' \
  --data '{"url": "https://gugudata.github.io/gugudata-io/guides/article-content-extraction-api-seo-guide/"}'
View sample Demo response

This is a fixed Demo sample, shortened for readability. It is not the response to your example request. Your results depend on the input and dataset. File URLs are temporary and have been removed.

{
  "dataStatus": {
    "statusCode": 200,
    "status": "SUCCESS",
    "statusDescription": "successfully",
    "dataTotalCount": 1
  },
  "data": {
    "url": "https://gugudata.github.io/gugudata-io/guides/article-content-extraction-api-seo-guide/",
    "title": "Article Content Extraction API Integration Guide | GuGuData.io Guides",
    "description": "Extract LLM-ready article text, readable HTML, metadata, images, author, publication time, and source for RAG, search, monitoring, and analysis.",
    "content": "<div><div><h1>Article Content Extraction API Integration Guide</h1>\n<p>Search teams, content platforms, and data products often start with the same messy input: a public article URL. The page may include navigation, cookie banners, related posts, comments, scripts, and layout HTML. What you usually need for an SEO workflow is much smaller and much more structured: the article title, description, readable body, plain text, main image, image candidates, author, published time, and source domain.</p>\n<p>The <a href=\"https://gugudata.io/details/fetchcontent\">GuGuData Article Content Extraction API</a> is designed for that exact job. It extracts clean article content from a public webpage URL and returns a normalized JSON response that can be stored, indexed, summarized, compared, or passed int…",
    "contentText": "Article Content Extraction API Integration Guide\nSearch teams, content platforms, and data products often start with the same messy input: a public article URL. The page may include navigation, cookie banners, related posts, comments, scripts, and layout HTML. What you usually need for an SEO workflow is much smaller and much more structured: the article title, description, readable body, plain text, main image, image candidates, author, published time, and source domain.\nThe GuGuData Article Content Extraction API is designed for that exact job. It extracts clean article content from a public webpage URL and returns a normalized JSON response that can be stored, indexed, summarized, compared, or passed into downstream SEO automation.\nThis guide explains where the API fits in an SEO stack,…",
    "image": "",
    "images": [],
    "author": "",
    "published": "2026-07-08 00:00",
    "source": "gugudata.github.io"
  }
}

Limits and things to check

These APIs do not bypass logins, paywalls, or source restrictions. Extraction depends on page structure; author, publication time, and images may be absent. Validate output and follow the source site’s permitted uses.

Normal paid API subscriptions include unlimited total requests during the active paid term, with up to 5 requests per second. Free trials have a limited allowance: sign in and check Trial Keys for the balance. Read each product page for pricing, renewal, and cancellation details.