API features
Scanned PDF recognition
Upload a PDF and receive page-level recognized text plus combined full text for search, review, and automation.
Page-by-page text
Receive one text entry per page in document order, including empty entries for blank pages.
Combined document text
Store extracted text beside the source PDF so historical scans, contracts, receipts, and forms become searchable.
English and Chinese options
Choose English, Simplified Chinese, Traditional Chinese, or an English and Chinese combination.
Document workflows
Call the endpoint from document queues, ingestion services, or back-office automation for repeated PDF processing.
Upload or PDF link
Upload a PDF or submit a public PDF link for automated document intake.
API Document
HTTP Protocol:HTTPS
HTTP Method:POST
HTTP Endpoint:https://api.gugudata.io/v1/imagerecognition/pdf2text
Response Type:application/json; charset=utf-8
DEMO Endpoint:https://api.gugudata.io/v1/imagerecognition/pdf2text/demo
Live Demo:Try Interactive Demo
Postman Collection:Run in Postman
API Request Parameters
| Name | Type | Is Required | Default Value | Remark |
|---|---|---|---|---|
| appkey | string | true | YOUR_APPKEY | API key used to authenticate the request. |
| file | file | false | Upload a PDF. Provide exactly one of file or file_url. Maximum 20 MiB and 50 pages. | |
| file_url | string | false | Public HTTP(S) PDF URL, up to 2048 characters. Provide exactly one of file or file_url. | |
| lang | string | false | en | en, zh-CN, zh-TW, en+zh-CN, or en+zh-TW. Recognition quality depends on the scan. |
API Response Parameters
| Name | Type | Remark |
|---|---|---|
| resultText | array<string> | Recognized text in page order; blank pages retain an empty string. |
| text | string | Nonempty page text joined by two newline characters. |
API Response Status Codes
| Status Code | Explanation of Status Code | Remarks |
|---|---|---|
| 200 | Request processed successfully. | Some endpoints expose a separate application-level status field in the response body, such as `dataStatus.statusCode`. |
| 400 | Invalid request parameters or request format. | Check required fields, data types, and request body format. |
| 401 | Missing or unknown application key. | Provide a valid `appkey` with the request. |
| 403 | The application key is recognized but access is not allowed. | The key may be expired, inactive, or not permitted for the requested API. |
| 413 | PDF exceeds the size or page limit. | Use a PDF of at most 20 MiB and 50 pages. |
| 422 | PDF cannot be recognized. | Use an unencrypted PDF with legible printed text. |
| 429 | Request rate or trial usage limit exceeded. | Reduce request frequency and retry with backoff. If a trial allowance is exhausted, check Trial Keys for remaining requests and subscribe to continue. |
| 500 | Internal service error. | Retry later or contact support if the error persists. |
| 502 | PDF download or recognition failed. | Check the PDF link and try again later. |
| 503 | Recognition is temporarily unavailable. | Try again later. |
Remote MCP
Use this API through GuGuData Remote MCP
Connect your AI client once, authorize in the browser, and the client can use the GuGuData API tools available to your account. You do not need to paste an appkey into the MCP client.
https://mcp.gugudata.io/mcp{
"mcpServers": {
"gugudata": {
"url": "https://mcp.gugudata.io/mcp",
"transportType": "streamable-http"
}
}
}Code Snippets
PDF to Text
Extract text from a PDF for review, search, or a downstream document workflow.
Input and coverage: Maximum 20 MiB and 50 pages. Supply either file or file_url. Supported lang values are en, zh-CN, zh-TW, en+zh-CN, and en+zh-TW.
Example request
Replace YOUR_APPKEY with your own API key. This request uses the normal API endpoint.
curl --fail-with-body 'https://api.gugudata.io/v1/imagerecognition/pdf2text?appkey=YOUR_APPKEY' \
--form 'file_url=https://gugudata.io/assets/demo/pdf2text-demo.pdf' \
--form 'lang=en'View sample Demo response
This is a fixed Demo sample, shortened for readability. It is not the response to your example request. Your results depend on the input and dataset. File URLs are temporary and have been removed.
{
"dataStatus": {
"statusCode": 200,
"status": "SUCCESS",
"statusDescription": "successfully",
"dataTotalCount": 1
},
"data": {
"resultText": [
"GuGuData PDF OCR demo"
],
"text": "GuGuData PDF OCR demo\n\nSecond page document workflow"
}
}Frequently asked questions
Common questions about request limits, trials, renewal, and cancellation.
What does the 5 requests per second limit mean?
It is a request-rate limit, not a concurrency allowance. Keep your request rate at or below 5 per second. If the service returns a rate-limit response, reduce the frequency and retry with backoff.
Is total paid usage limited?
There is no total request cap during an active paid subscription and no usage-based overage charge.
How does the free trial work?
Each account can receive a limited trial key. Requests are shared across APIs in the same category. Check Trial Keys for the exact allowance and remaining requests.
How do renewal and cancellation work?
The plan renews annually until canceled. Manage or cancel the renewal from Dashboard → Orders & Billing. Access continues through the paid period after cancellation. See the Terms of Service for the current refund policy. If the billing portal is unavailable, contact support.
Service terms
Read our service terms and privacy policy before subscribing.




