Webpage Readable Content Extraction

Extract clean, readable article content from webpages or HTML

Intelligent Extraction Various Elements Information

API features

Readable article HTML

Extract the main article with useful text formatting.

Plain-text output

Use extracted text for indexing and content workflows.

Article metadata

Read the title, author and language when available.

Page or HTML input

Submit a public page URL or raw HTML.

Reduced page clutter

Remove navigation and surrounding page elements.

Content workflows

Prepare articles for search, review and summarization.

API Document

HTTP Protocol:HTTPS

HTTP Method:POST

HTTP Endpoint:https://api.gugudata.io/v1/websitetools/readability

Response Type:application/json; charset=utf-8

DEMO Endpoint:https://api.gugudata.io/v1/websitetools/readability/demo

Live Demo:Try Interactive Demo

Postman Collection:Run in Postman

API Request Parameters

NameTypeIs RequiredDefault ValueRemark
appkey string trueYOUR_APPKEY API key used to authenticate the request.
html string falseYOUR_VALUE Raw HTML content. Supply either `html` or `url`.
url string falseYOUR_VALUE Target webpage URL. Supply either `url` or `html`.

API Response Parameters

NameTypeRemark
DataStatus.RequestParameter string Privacy-safe summary of the normalized request.
DataStatus.StatusCode integer Application-level status code.
DataStatus.StatusDescription string Human-readable application status message.
DataStatus.ResponseDateTime string Time at which the response was generated.
DataStatus.DataTotalCount integer Number of matching or returned records for this request.
Data.Title string Article title.
Data.Byline string Article author.
Data.Dir string Article text direction.
Data.Lang string Article language.
Data.Content string Article content.
Data.TextContent string Article content (without HTML tags, divided by paragraphs).
Data.Length integer Article length.
Data.Excerpt string Article excerpt.
Data.SiteName string Website name.
Data.PublishedTime array<string> Article publication time.

API Response Status Codes

Status CodeExplanation of Status CodeRemarks
200Request processed successfully.Some endpoints expose a separate application-level status field in the response body, such as `dataStatus.statusCode`.
400Invalid request parameters or request format.Check required fields, data types, and request body format.
401Missing or unknown application key.Provide a valid `appkey` with the request.
403The application key is recognized but access is not allowed.The key may be expired, inactive, or not permitted for the requested API.
429Request rate or trial usage limit exceeded.Reduce request frequency and retry with backoff. If a trial allowance is exhausted, check Trial Keys for remaining requests and subscribe to continue.
500Internal service error.Retry later or contact support if the error persists.
503Upstream service unavailable.Retry later; the requested upstream dependency is temporarily unavailable.

Remote MCP

AI client access

Use this API through GuGuData Remote MCP

Connect your AI client once, authorize in the browser, and the client can use the GuGuData API tools available to your account. You do not need to paste an appkey into the MCP client.

Endpoint https://mcp.gugudata.io/mcp
Open MCP Guide
Generic MCP configuration
{
  "mcpServers": {
    "gugudata": {
      "url": "https://mcp.gugudata.io/mcp",
      "transportType": "streamable-http"
    }
  }
}

Code Snippets

Readable Article Extraction

Submit a public article URL or HTML and extract readable HTML, text, and available metadata.

Input and coverage: Supply exactly one of url or html. URL maximum 2048 characters; HTML maximum 1 MiB in UTF-8. Some pages have no extractable article.

Example request

Replace YOUR_APPKEY with your own API key. This request uses the normal API endpoint.

curl --fail-with-body 'https://api.gugudata.io/v1/websitetools/readability?appkey=YOUR_APPKEY' \
  --request POST \
  --header 'Content-Type: application/json' \
  --data '{"url": "https://gugudata.github.io/gugudata-io/guides/article-content-extraction-api-seo-guide/"}'
View sample Demo response

This is a fixed Demo sample, shortened for readability. It is not the response to your example request. Your results depend on the input and dataset. File URLs are temporary and have been removed.

{
  "dataStatus": {
    "statusCode": 200,
    "status": "SUCCESS",
    "statusDescription": "successfully",
    "dataTotalCount": 1
  },
  "data": {
    "title": "Article Content Extraction API Integration Guide",
    "byline": "GuGuData.io",
    "dir": null,
    "lang": "en",
    "content": "<div id=\"readability-page-1\" class=\"page\"><div><h2 id=\"article-content-extraction-api-integration-guide\">Article Content Extraction API Integration Guide</h2>\n<p>Search teams, content platforms, and data products often start with the same messy input: a public article URL. The page may include navigation, cookie banners, related posts, comments, scripts, and layout HTML. What you usually need for an SEO workflow is much smaller and much more structured: the article title, description, readable body, plain text, main image, image candidates, author, published time, and source domain.</p>\n<p>The <a href=\"https://gugudata.io/details/fetchcontent\">GuGuData Article Content Extraction API</a> is designed for that exact job. It extracts clean article content from a public webpage URL and returns …",
    "textContent": "Article Content Extraction API Integration Guide\nSearch teams, content platforms, and data products often start with the same messy input: a public article URL. The page may include navigation, cookie banners, related posts, comments, scripts, and layout HTML. What you usually need for an SEO workflow is much smaller and much more structured: the article title, description, readable body, plain text, main image, image candidates, author, published time, and source domain.\nThe GuGuData Article Content Extraction API is designed for that exact job. It extracts clean article content from a public webpage URL and returns a normalized JSON response that can be stored, indexed, summarized, compared, or passed into downstream SEO automation.\nThis guide explains where the API fits in an SEO stack,…",
    "length": 10436,
    "excerpt": "Extract LLM-ready article text, readable HTML, metadata, images, author, publication time, and source for RAG, search, monitoring, and analysis.",
    "siteName": "GuGuData.io",
    "publishedTime": "2026-07-08T00:00:00.000Z"
  }
}

Frequently asked questions

Common questions about request limits, trials, renewal, and cancellation.

What does the 5 requests per second limit mean?

It is a request-rate limit, not a concurrency allowance. Keep your request rate at or below 5 per second. If the service returns a rate-limit response, reduce the frequency and retry with backoff.

Is total paid usage limited?

There is no total request cap during an active paid subscription and no usage-based overage charge.

How does the free trial work?

Each account can receive a limited trial key. Requests are shared across APIs in the same category. Check Trial Keys for the exact allowance and remaining requests.

How do renewal and cancellation work?

The plan renews annually until canceled. Manage or cancel the renewal from Dashboard → Orders & Billing. Access continues through the paid period after cancellation. See the Terms of Service for the current refund policy. If the billing portal is unavailable, contact support.

Technical support

Need help choosing an API or getting started?

[email protected]

Service terms

Read our service terms and privacy policy before subscribing.