Bestseller
Article Extraction Content Parsing Web Content

Article Extractor

Extract structured article content and metadata from URLs or HTML

What can this API do?

Public URL or raw HTML input

Extract article content from a live public webpage or from HTML supplied directly by your application.

Structured editorial metadata

Receive the title, description, author, publication date, source, favicon, and article type in one record.

Readable article body

Isolate the primary article content from surrounding navigation, promotional blocks, and page furniture.

Links and image context

Capture page links, the lead image, image alternatives, and related visual context when available.

Estimated reading time

Use the returned reading-time value to support previews, prioritization, and editorial planning.

Research-ready content records

Standardize articles for monitoring, archiving, content analysis, and downstream knowledge workflows.

Article Extractor
Annual plan
$49$99
Try it for free!
Sign in to get a trial key and test all APIs.
Sign In
Secure checkout powered by Stripe

API Document

HTTP Protocol:HTTPS

HTTP Method:POST

HTTP Endpoint:https://api.gugudata.io/v1/article/extract

Response Type:application/json; charset=utf-8

DEMO Endpoint:https://api.gugudata.io/v1/article/extract/demo

Live Demo:Try Interactive Demo

Postman Collection:Run in Postman

API Request Parameters

NameTypeIs RequiredDefault ValueRemark
appkeystringtrueYOUR_APPKEY API key used to authenticate the request.
urlstringtrue Target webpage URL.

API Response Parameters

NameTypeRemark
DataStatus.StatusCodeinteger Application-level status code.
DataStatus.StatusDescriptionstring Human-readable application status message.
DataStatus.ResponseDateTimestring Time at which the response was generated.
DataStatus.DataTotalCountinteger Number of matching or returned records for this request.
Data.urlstring Source URL of the article.
Data.titlestring Extracted article title.
Data.descriptionstring Article description/summary.
Data.linksarray<string> Array of links contained in the article.
Data.imagestring Main article image URL.
Data.contentstring Extracted article content (HTML format, with ads and navigation removed).
Data.authorstring Article author (if available, may be empty string).
Data.faviconstring Website favicon URL.
Data.sourcestring Source website domain (e.g., sohu.com).
Data.publishedstring Article publication date/time (format: YYYY-MM-DD HH:MM).
Data.ttrinteger Estimated reading time (Time to Read, in minutes).
Data.typestring Article type (e.g., news, article, etc.).

API Response Status Codes

Status CodeExplanation of Status CodeRemarks
200Request processed successfully.Some endpoints expose a separate application-level status field in the response body, such as `dataStatus.statusCode`.
400Invalid request parameters or request format.Check required fields, data types, and request body format.
401Missing or unknown application key.Provide a valid `appkey` with the request.
403The application key is recognized but access is not allowed.The key may be expired, inactive, or not permitted for the requested API.
429Request rate or trial usage limit exceeded.Reduce concurrency or retry after the limit window resets.
500Internal service error.Retry later or contact support if the error persists.
503Upstream service unavailable.Retry later; the requested upstream dependency is temporarily unavailable.

Remote MCP

AI client access

Use this API through GuGuData Remote MCP

Connect your AI client once, authorize in the browser, and the client can use the GuGuData API tools available to your account. You do not need to paste an appkey into the MCP client.

Endpoint https://mcp.gugudata.io/mcp
Open MCP Guide
Generic MCP configuration
{
  "mcpServers": {
    "gugudata": {
      "url": "https://mcp.gugudata.io/mcp",
      "transportType": "streamable-http"
    }
  }
}

Code Snippets