
LLM-ready Article Content Extraction API
Extract clean, LLM-ready article content and metadata
Extract structured article content and metadata from URLs or HTML
Extract article content from a live public webpage or from HTML supplied directly by your application.
Receive the title, description, author, publication date, source, favicon, and article type in one record.
Isolate the primary article content from surrounding navigation, promotional blocks, and page furniture.
Capture page links, the lead image, image alternatives, and related visual context when available.
Use the returned reading-time value to support previews, prioritization, and editorial planning.
Standardize articles for monitoring, archiving, content analysis, and downstream knowledge workflows.

HTTP Protocol:HTTPS
HTTP Method:POST
HTTP Endpoint:https://api.gugudata.io/v1/article/extract
Response Type:application/json; charset=utf-8
DEMO Endpoint:https://api.gugudata.io/v1/article/extract/demo
Live Demo:Try Interactive Demo
Postman Collection:Run in Postman
| Name | Type | Is Required | Default Value | Remark |
|---|---|---|---|---|
| appkey | string | true | YOUR_APPKEY | API key used to authenticate the request. |
| url | string | true | Target webpage URL. |
| Name | Type | Remark |
|---|---|---|
| DataStatus.StatusCode | integer | Application-level status code. |
| DataStatus.StatusDescription | string | Human-readable application status message. |
| DataStatus.ResponseDateTime | string | Time at which the response was generated. |
| DataStatus.DataTotalCount | integer | Number of matching or returned records for this request. |
| Data.url | string | Source URL of the article. |
| Data.title | string | Extracted article title. |
| Data.description | string | Article description/summary. |
| Data.links | array<string> | Array of links contained in the article. |
| Data.image | string | Main article image URL. |
| Data.content | string | Extracted article content (HTML format, with ads and navigation removed). |
| Data.author | string | Article author (if available, may be empty string). |
| Data.favicon | string | Website favicon URL. |
| Data.source | string | Source website domain (e.g., sohu.com). |
| Data.published | string | Article publication date/time (format: YYYY-MM-DD HH:MM). |
| Data.ttr | integer | Estimated reading time (Time to Read, in minutes). |
| Data.type | string | Article type (e.g., news, article, etc.). |
| Status Code | Explanation of Status Code | Remarks |
|---|---|---|
| 200 | Request processed successfully. | Some endpoints expose a separate application-level status field in the response body, such as `dataStatus.statusCode`. |
| 400 | Invalid request parameters or request format. | Check required fields, data types, and request body format. |
| 401 | Missing or unknown application key. | Provide a valid `appkey` with the request. |
| 403 | The application key is recognized but access is not allowed. | The key may be expired, inactive, or not permitted for the requested API. |
| 429 | Request rate or trial usage limit exceeded. | Reduce concurrency or retry after the limit window resets. |
| 500 | Internal service error. | Retry later or contact support if the error persists. |
| 503 | Upstream service unavailable. | Retry later; the requested upstream dependency is temporarily unavailable. |
Connect your AI client once, authorize in the browser, and the client can use the GuGuData API tools available to your account. You do not need to paste an appkey into the MCP client.
https://mcp.gugudata.io/mcp{
"mcpServers": {
"gugudata": {
"url": "https://mcp.gugudata.io/mcp",
"transportType": "streamable-http"
}
}
}
Extract clean, LLM-ready article content and metadata

Extract ordered image candidates from a public article

Extract links and destinations from a public webpage

Extract clean, readable article content from webpages or HTML