API features
Readable HTML and clean text
Receive the article body as cleaned HTML and paragraph-preserving plain text for different downstream needs.
Complete article metadata
Capture the final URL, title, description, author, publication time, and source in one request.
Primary image and image candidates
Use the main image plus ordered, deduplicated image records with alt text and available dimensions.
LLM and RAG-ready output
Feed normalized article text into retrieval, summarization, classification, and knowledge workflows.
Multilingual webpage support
Extract readable content from common international encodings, including English and Chinese pages.
Search and monitoring workflows
Standardize public articles for indexing, competitive research, archives, and scheduled content analysis.
API Document
HTTP Protocol:HTTPS
HTTP Method:POST
HTTP Endpoint:https://api.gugudata.io/v1/websitetools/fetchcontent
Response Type:application/json; charset=utf-8
DEMO Endpoint:https://api.gugudata.io/v1/websitetools/fetchcontent/demo
Live Demo:Try Interactive Demo
Postman Collection:Run in Postman
API Request Parameters
| Name | Type | Is Required | Default Value | Remark |
|---|---|---|---|---|
| appkey | string | true | YOUR_APPKEY | API key used to authenticate the request. |
| url | string | true | https://gugudata.github.io/gugudata-io/guides/article-content-extraction-api-seo-guide/ | Complete public HTTP or HTTPS article URL. |
API Response Parameters
| Name | Type | Remark |
|---|---|---|
| url | string | Final article URL after redirects. |
| title | string | Extracted article title or an empty string. |
| description | string | Article summary or an empty string. |
| content | string | Clean readable article HTML. |
| contentText | string | Plain article text with paragraph boundaries. |
| image | string | Primary article image URL or an empty string. |
| images | array<object> | Stable image list with src, absoluteUrl, alt, width, and height. |
| author | string | Article author or an empty string. |
| published | string | Publication time or an empty string. |
| source | string | Source publisher or host. |
API Response Status Codes
| Status Code | Explanation of Status Code | Remarks |
|---|---|---|
| 200 | Request processed successfully. | Some endpoints expose a separate application-level status field in the response body, such as `dataStatus.statusCode`. |
| 400 | Invalid request parameters or request format. | Check required fields, data types, and request body format. |
| 401 | Missing or unknown application key. | Provide a valid `appkey` with the request. |
| 403 | The application key is recognized but access is not allowed. | The key may be expired, inactive, or not permitted for the requested API. |
| 422 | The target response cannot be extracted as an article. | Use an HTML article page with readable body content and a valid redirect chain. |
| 429 | Request rate or trial usage limit exceeded. | Reduce request frequency and retry with backoff. If a trial allowance is exhausted, check Trial Keys for remaining requests and subscribe to continue. |
| 502 | The target website could not be reached successfully. | Retry later or verify that the target website is available. |
| 503 | The extraction service is temporarily unavailable. | Retry later or contact support if the error persists. |
Remote MCP
Use this API through GuGuData Remote MCP
Connect your AI client once, authorize in the browser, and the client can use the GuGuData API tools available to your account. You do not need to paste an appkey into the MCP client.
https://mcp.gugudata.io/mcp{
"mcpServers": {
"gugudata": {
"url": "https://mcp.gugudata.io/mcp",
"transportType": "streamable-http"
}
}
}Code Snippets
Webpage Content Extraction
Extract article text and metadata from a public webpage for search or content processing.
Input and coverage: Use a public HTTP(S) article URL. Page access restrictions and source structure can affect extraction; metadata may be absent.
Example request
Replace YOUR_APPKEY with your own API key. This request uses the normal API endpoint.
curl --fail-with-body 'https://api.gugudata.io/v1/websitetools/fetchcontent?appkey=YOUR_APPKEY' \
--request POST \
--header 'Content-Type: application/json' \
--data '{"url": "https://gugudata.github.io/gugudata-io/guides/article-content-extraction-api-seo-guide/"}'View sample Demo response
This is a fixed Demo sample, shortened for readability. It is not the response to your example request. Your results depend on the input and dataset. File URLs are temporary and have been removed.
{
"dataStatus": {
"statusCode": 200,
"status": "SUCCESS",
"statusDescription": "successfully",
"dataTotalCount": 1
},
"data": {
"url": "https://gugudata.github.io/gugudata-io/guides/article-content-extraction-api-seo-guide/",
"title": "Article Content Extraction API Integration Guide | GuGuData.io Guides",
"description": "Extract LLM-ready article text, readable HTML, metadata, images, author, publication time, and source for RAG, search, monitoring, and analysis.",
"content": "<div><div><h1>Article Content Extraction API Integration Guide</h1>\n<p>Search teams, content platforms, and data products often start with the same messy input: a public article URL. The page may include navigation, cookie banners, related posts, comments, scripts, and layout HTML. What you usually need for an SEO workflow is much smaller and much more structured: the article title, description, readable body, plain text, main image, image candidates, author, published time, and source domain.</p>\n<p>The <a href=\"https://gugudata.io/details/fetchcontent\">GuGuData Article Content Extraction API</a> is designed for that exact job. It extracts clean article content from a public webpage URL and returns a normalized JSON response that can be stored, indexed, summarized, compared, or passed int…",
"contentText": "Article Content Extraction API Integration Guide\nSearch teams, content platforms, and data products often start with the same messy input: a public article URL. The page may include navigation, cookie banners, related posts, comments, scripts, and layout HTML. What you usually need for an SEO workflow is much smaller and much more structured: the article title, description, readable body, plain text, main image, image candidates, author, published time, and source domain.\nThe GuGuData Article Content Extraction API is designed for that exact job. It extracts clean article content from a public webpage URL and returns a normalized JSON response that can be stored, indexed, summarized, compared, or passed into downstream SEO automation.\nThis guide explains where the API fits in an SEO stack,…",
"image": "",
"images": [],
"author": "",
"published": "2026-07-08 00:00",
"source": "gugudata.github.io"
}
}Frequently asked questions
Common questions about request limits, trials, renewal, and cancellation.
What does the 5 requests per second limit mean?
It is a request-rate limit, not a concurrency allowance. Keep your request rate at or below 5 per second. If the service returns a rate-limit response, reduce the frequency and retry with backoff.
Is total paid usage limited?
There is no total request cap during an active paid subscription and no usage-based overage charge.
How does the free trial work?
Each account can receive a limited trial key. Requests are shared across APIs in the same category. Check Trial Keys for the exact allowance and remaining requests.
How do renewal and cancellation work?
The plan renews annually until canceled. Manage or cancel the renewal from Dashboard → Orders & Billing. Access continues through the paid period after cancellation. See the Terms of Service for the current refund policy. If the billing portal is unavailable, contact support.
Service terms
Read our service terms and privacy policy before subscribing.




