Skip to main content
POST
/api/v1/augment/text-parser
Dies ist eine experimentelle API. Das Request- und Response-Format kann sich ohne Vorankündigung ändern.
Laden Sie eine Dokumentdatei per multipart/form-data über das Feld file hoch, bis zu 25 MB. Die unterstützten Formate gehen über die Menge strukturierter Dokumente hinaus. Der Endpunkt akzeptiert:
  • Dokumente — PDF, EPUB, DOCX, PPTX, XLSX, XLS
  • Klartext und Daten — Klartext, Markdown, CSV, JSON, YAML, TOML, XML und jeden text/*-MIME-Typ
  • Quellcode — rund 140 erkannte Dateiendungen, darunter .py, .ts, .js, .go, .rs, .c, .cpp, .java, .ps1, .sh, .sql, sowie Dateien ohne Endung wie Dockerfile und Makefile
Wenn Ihre Eingabe bereits Text oder Code ist, besteht keine Notwendigkeit, sie zuerst in PDF zu konvertieren.
Nur die modernen Office-Formate werden unterstützt. Die alten Binärformate .doc und .ppt werden nicht akzeptiert — konvertieren Sie sie zuerst in .docx bzw. .pptx. Das alte .xls wird hingegen sehr wohl akzeptiert.
Setzen Sie response_format auf json (Standard) für strukturierte Ausgabe mit extrahiertem Text und Token-Anzahl, oder auf text für den reinen extrahierten Text. Datenschutz: Das Text-Parsing läuft vollständig im Arbeitsspeicher auf der Infrastruktur von Venice mit Zero Data Retention. Ihre Dokumente werden verarbeitet und sofort verworfen — kein Inhalt wird gespeichert oder protokolliert. Preise: $0.01 pro Anfrage.

Beispiel (cURL)


Autorisierungen

Authorization
string
header
erforderlich

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

multipart/form-data
file
file
erforderlich

The document file to parse. Supported formats: PDF, DOCX, PPTX, XLSX, and plain text files. Maximum size: 25MB.

response_format
enum<string>
Standard:json

The format of the response output. "json" returns structured JSON with text and token count, "text" returns only the extracted text.

Verfügbare Optionen:
json,
text

Antwort

Text extraction completed successfully

Text parser response containing extracted text and token count.

text
string
erforderlich

The extracted text content from the document.

tokens
number
erforderlich

The token count of the extracted text.