> ## Documentation Index
> Fetch the complete documentation index at: https://cortex-ad5578da.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Upload Content

> Upload raw text or markdown content for ingestion.
Supports both single and batch uploads via the contents list in the request.

export const TableOfContents = ({title = 'On this page', items, minHeadingLevel = 2, maxHeadingLevel = 3, className = '', activeClassName = 'text-primary dark:text-primary-light border-primary dark:border-primary-light hover:border-primary dark:hover:border-primary-light', inactiveClassName = 'hover:text-gray-900 dark:text-gray-400 dark:hover:text-gray-300'}) => {
  const [toc, setToc] = useState(items ?? []);
  const [activeId, setActiveId] = useState('');
  useEffect(() => {
    if (items && items.length) return;
    if (typeof document === 'undefined') return;
    const selectors = [];
    for (let lvl = minHeadingLevel; lvl <= maxHeadingLevel; lvl++) {
      selectors.push(`h${lvl}`);
    }
    const nodes = Array.from(document.querySelectorAll(selectors.join(','))).filter(el => el.id);
    const built = [];
    let currentTop = null;
    nodes.forEach(el => {
      const level = Number(el.tagName.slice(1));
      const node = {
        id: el.id,
        label: el.textContent.trim(),
        href: `#${el.id}`,
        children: []
      };
      if (level === minHeadingLevel) {
        built.push(node);
        currentTop = node;
      } else if (level > minHeadingLevel && currentTop) {
        currentTop.children.push(node);
      } else {
        built.push(node);
      }
    });
    setToc(built);
  }, [items, minHeadingLevel, maxHeadingLevel]);
  useEffect(() => {
    if (typeof document === 'undefined') return;
    const applyHash = () => {
      const h = window.location.hash.replace('#', '');
      if (h) setActiveId(h);
    };
    applyHash();
    const observer = new IntersectionObserver(entries => {
      entries.forEach(entry => {
        if (entry.isIntersecting) {
          setActiveId(entry.target.id);
        }
      });
    }, {
      rootMargin: '0px 0px -70% 0px',
      threshold: 0.1
    });
    const ids = toc.flatMap(i => [i, ...i.children ?? []]).map(i => i.id);
    ids.forEach(id => {
      const el = document.getElementById(id);
      if (el) observer.observe(el);
    });
    window.addEventListener('hashchange', applyHash);
    return () => {
      window.removeEventListener('hashchange', applyHash);
      observer.disconnect();
    };
  }, [toc]);
  const Item = ({node, depth = 0}) => {
    const isActive = activeId === node.id;
    return <li className="toc-item relative ml-6" data-depth={depth}>
        <a href={node.href} className={`py-1 block font-medium ${isActive ? activeClassName : inactiveClassName}`} style={depth > 0 ? {
      marginLeft: `${depth}rem`
    } : {}} onClick={() => setActiveId(node.id)}>
          {node.label}
        </a>
        {node.children && node.children.length > 0 && <>
            {node.children.map(child => <Item node={child} depth={depth + 1} key={child.href} />)}
          </>}
      </li>;
  };
  const data = toc && toc.length ? toc : items || [];
  return <div className={`text-gray-600 text-sm leading-6 w-[18rem] pb-4 -mt-10 pt-10 hidden xl:block ${className}`} id="table-of-contents-custom">
      <ul id="table-of-contents-custom-content" className="toc">
        <li className="toc-item relative">
          <div className="text-gray-700 dark:text-gray-300 font-medium flex items-center space-x-2 py-1">
            <svg width="16" height="16" viewBox="0 0 16 16" fill="none" stroke="currentColor" strokeWidth="2" xmlns="http://www.w3.org/2000/svg" className="h-3 w-3">
              <path d="M2.44434 12.6665H13.5554" strokeLinecap="round" strokeLinejoin="round"></path>
              <path d="M2.44434 3.3335H13.5554" strokeLinecap="round" strokeLinejoin="round"></path>
              <path d="M2.44434 8H7.33323" strokeLinecap="round" strokeLinejoin="round"></path>
            </svg>
            <span>{title}</span>
          </div>
        </li>
        {data.map(node => <Item node={node} key={node.href} />)}
      </ul>
    </div>;
};

<Panel>
  <TableOfContents />
</Panel>

<Tip> Hit the `Try it` button to try this API now in our playground. It's the best way to check the full request and response in one place, customize your parameters, and generate ready-to-use code snippets.</Tip>

### Examples

<Tabs>
  <Tab title="API Request">
    ```bash theme={null}
    curl -X 'POST' \
    'https://api.usecortex.ai/ingestion/upload-content' \
    -H 'accept: application/json' \
    -H 'Content-Type: application/json' \
    -d '{
    "content": {
    "contents": [
      {
        "file_id": "string",
        "content": "# Introduction\n\nThis is the document content.",
        "is_markdown": false,
        "tenant_metadata": "",
        "document_metadata": "",
        "relations": false
      }
    ]
    },
    "tenant_id": "string",
    "sub_tenant_id": "",
    "upsert": true
    }'
    ```
  </Tab>

  <Tab title="TypeScript">
    ```ts theme={null}
    const result = await client.upload.uploadText({
      tenant_id: "tenant_1234",
      sub_tenant_id: "sub_tenant_4567",
      body: {
        content: "Your text content here",
        file_id: "text_doc_123456",
        tenant_metadata: {},
        document_metadata: {}
      }
    });
    ```
  </Tab>

  <Tab title="Python (Sync)">
    ```python theme={null}
    # Async usage is similar, just use async_client and await
    result = client.upload.upload_text(
        tenant_id="tenant_1234",
        sub_tenant_id="sub_tenant_4567",
        content="Your text content here",
        file_id="text_doc_123456",
        tenant_metadata={},
        document_metadata={}
    )
    ```
  </Tab>
</Tabs>

Upload text content directly to your tenant's knowledge base. The text will be processed, chunked, and indexed for search and retrieval.

## Text Processing Pipeline

When you upload text content, it goes through a streamlined processing pipeline optimized for direct text input:

### 1. **Immediate Upload & Queue**

* Your text content is immediately accepted and stored securely
* It's added to our processing queue for background processing
* You receive a confirmation response with a `file_id` for tracking

### 2. **Text Processing Phase**

Our system automatically handles:

* **Content Validation**: Ensuring text content is properly formatted and accessible
* **Format Detection**: Identifying markdown, plain text, or structured content
* **Text Normalization**: Cleaning and standardizing text formatting

### 3. **Intelligent Chunking**

* Text is split into semantically meaningful chunks
* Chunk size is optimized for both context preservation and search accuracy
* Overlapping boundaries ensure no information is lost between chunks
* Metadata is preserved and associated with each chunk

### 4. **Embedding Generation**

* Each chunk is converted into high-dimensional vector embeddings
* Embeddings capture semantic meaning and context
* Vectors are optimized for similarity search and retrieval

### 5. **Indexing & Database Updates**

* Embeddings are stored in our vector database for fast similarity search
* Full-text search indexes are created for keyword-based queries
* Metadata is indexed for filtering and faceted search
* Cross-references are established for related content

### 6. **Quality Assurance**

* Automated quality checks ensure processing accuracy
* Content validation verifies text completeness
* Embedding quality is assessed for optimal retrieval performance

<Note>
  **Processing Time**: Text content is typically processed and searchable within 1-3 minutes. Large text blocks (10,000+ words) may take up to 5 minutes. You can check processing status using the document ID returned in the response.
</Note>

<Note>
  **Default Sub-Tenant Behavior**: If you don't specify a `sub_tenant_id`, the text content will be uploaded to the default sub-tenant created when your tenant was set up. This is perfect for organization-wide content that should be accessible across all departments.
</Note>

> **File ID Management**: The system uses a priority-based approach for file ID assignment:
>
> 1. **First Priority**: If you provide a `file_id` as a direct body parameter, that specific ID will be used
> 2. **Second Priority**: If no direct `file_id` is provided, the system checks for a `file_id` in the `document_metadata` object
> 3. **Auto-Generation**: If neither source provides a `file_id`, the system will automatically generate a unique identifier

### **Duplicate File ID Behavior**

When you upload text content with a `file_id` that already exists in your tenant:

* **Overwrite Behavior**: The existing text content with the same `file_id` will be **completely replaced** with the new content
* **Processing**: The new text content will go through the full processing pipeline (validation, chunking, embedding generation, indexing)
* **Search Results**: Previous search results and embeddings from the old content will be replaced with the new content
* **Idempotency**: Uploading the same text content with the same `file_id` multiple times is safe and will result in the same final state

<Warning>
  **Important**: When overwriting existing text content, all previous chunks, embeddings, and search indexes associated with that `file_id` will be permanently removed and replaced. This action cannot be undone.
</Warning>

**Example Success Response for Duplicate File ID:**

```json theme={null}
{
  "message": "Text content uploaded successfully. Existing content with file_id 'text_123456' has been overwritten.",
  "file_id": "text_123456",
  "status": "success"
}
```

## Processing Status & Monitoring

After uploading, you can monitor your text content's processing status:

### **Immediate Response**

Upon successful upload, you'll receive:

```json theme={null}
{
  "message": "Text content uploaded successfully",
  "file_id": "doc_123456"
}
```

### **Processing States**

Your text content will progress through these states:

* **`queued`**: Text content is in the processing queue, waiting to be processed
* **`in_progress`**: Text content is actively being processed (includes validation, chunking, embedding generation, and indexing)
* **`success`**: Text content is fully processed and searchable
* **`errored`**: Processing encountered an error (rare occurrence)

<Info>
  **In-Progress Details**: While the status shows `in_progress`, the system is actually performing multiple steps: content validation, format detection, intelligent chunking, embedding generation, and database indexing. These happen sequentially but are all part of the single `in_progress` state.
</Info>

### **When Your Text is Ready**

Once processing is complete, your text content will be:

* ✅ **Searchable** via semantic search and Q\&A endpoints
* ✅ **Retrievable** through our retrieval APIs
* ✅ **Available** for AI-powered applications
* ✅ **Indexed** for fast query performance

<Warning>
  **Important**: Don't attempt to search or retrieve your text content immediately after upload. Wait for processing to complete (typically 1-3 minutes) to ensure optimal results.
</Warning>

## Error Responses

All endpoints return consistent error responses following the standard format. For detailed error information, see our [Error Responses](/api-reference/error-responses) documentation.


## OpenAPI

````yaml POST /ingestion/upload-content
openapi: 3.1.0
info:
  title: Cortex SDK API
  description: REST APIs for Cortex AI retrieval engine
  version: 0.0.1
servers:
  - url: /
    description: Local
    x-fern-server-name: cortex-backend-local
  - url: https://api.usecortex.ai
    description: Production
    x-fern-server-name: cortex-prod
    x-fern-audiences:
      - public
  - url: https://preprod.usecortex.ai
    description: Staging
    x-fern-server-name: cortex-staging
security: []
paths:
  /ingestion/upload-content:
    post:
      tags:
        - ingestion
      summary: Upload Content
      description: >-
        Upload raw text or markdown content for ingestion.

        Supports both single and batch uploads via the contents list in the
        request.
      operationId: upload_content_ingestion_upload_content_post
      requestBody:
        content:
          application/json:
            schema:
              $ref: >-
                #/components/schemas/Body_upload_content_ingestion_upload_content_post
        required: true
      responses:
        '200':
          description: Successful Response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/SourceUploadResponse'
        '400':
          description: Bad Request - Invalid input parameters
          content:
            application/json:
              schema:
                $ref: >-
                  #/components/schemas/cortex__models__response__commons__ActualErrorResponse
        '401':
          description: Unauthorized - Authentication required
          content:
            application/json:
              schema:
                $ref: >-
                  #/components/schemas/cortex__models__response__commons__ActualErrorResponse
        '403':
          description: Forbidden - Access denied
          content:
            application/json:
              schema:
                $ref: >-
                  #/components/schemas/cortex__models__response__commons__ActualErrorResponse
        '404':
          description: Not Found - Resource does not exist
          content:
            application/json:
              schema:
                $ref: >-
                  #/components/schemas/cortex__models__response__commons__ActualErrorResponse
        '422':
          description: Unprocessable Entity - Validation failed
          content:
            application/json:
              schema:
                $ref: >-
                  #/components/schemas/cortex__models__response__commons__ActualErrorResponse
        '500':
          description: Internal Server Error
          content:
            application/json:
              schema:
                $ref: >-
                  #/components/schemas/cortex__models__response__commons__ActualErrorResponse
        '503':
          description: Service Unavailable
          content:
            application/json:
              schema:
                $ref: >-
                  #/components/schemas/cortex__models__response__commons__ActualErrorResponse
      security:
        - HTTPBearer: []
components:
  schemas:
    Body_upload_content_ingestion_upload_content_post:
      properties:
        content:
          $ref: '#/components/schemas/ContentUploadRequest'
          description: CONTENT_DESCRIPTION
        tenant_id:
          type: string
          title: Tenant Id
          description: Unique identifier for the tenant/organization
          example: tenant_1234
        sub_tenant_id:
          type: string
          title: Sub Tenant Id
          description: >-
            Optional sub-tenant identifier used to organize data within a
            tenant. If omitted, the default sub-tenant created during tenant
            setup will be used.
          default: ''
          example: sub_tenant_4567
        upsert:
          type: boolean
          title: Upsert
          description: >-
            If true, update existing sources with the same source_id. Defaults
            to True.
          default: true
          example: true
      type: object
      required:
        - content
        - tenant_id
      title: Body_upload_content_ingestion_upload_content_post
    SourceUploadResponse:
      properties:
        success:
          type: boolean
          title: Success
          default: true
          example: true
        message:
          type: string
          title: Message
          default: Upload initiated successfully
        results:
          items:
            $ref: '#/components/schemas/SourceUploadResultItem'
          type: array
          title: Results
          description: List of upload results for each source.
          example: []
        success_count:
          type: integer
          title: Success Count
          description: Number of sources successfully queued.
          default: 0
          example: 1
        failed_count:
          type: integer
          title: Failed Count
          description: Number of sources that failed to upload.
          default: 0
          example: 1
      type: object
      title: SourceUploadResponse
    cortex__models__response__commons__ActualErrorResponse:
      properties:
        detail:
          $ref: >-
            #/components/schemas/cortex__models__response__commons__ErrorResponse
      type: object
      required:
        - detail
      title: ActualErrorResponse
    ContentUploadRequest:
      properties:
        contents:
          items:
            $ref: '#/components/schemas/ContentUploadItem'
          type: array
          title: Contents
          description: List of text/markdown content to upload.
          example: []
      type: object
      title: ContentUploadRequest
    SourceUploadResultItem:
      properties:
        source_id:
          type: string
          title: Source Id
          description: Unique identifier for the uploaded source.
          example: CortexDoc1234
        filename:
          anyOf:
            - type: string
            - type: 'null'
          title: Filename
          description: Original filename if present.
        status:
          $ref: '#/components/schemas/SourceStatus'
          description: Initial processing status.
          default: queued
        error:
          anyOf:
            - type: string
            - type: 'null'
          title: Error
          description: Error message if upload failed.
      type: object
      required:
        - source_id
      title: SourceUploadResultItem
    cortex__models__response__commons__ErrorResponse:
      properties:
        success:
          type: boolean
          title: Success
          default: false
          example: true
        message:
          type: string
          title: Message
          default: Error occurred
        error_code:
          anyOf:
            - type: string
            - type: 'null'
          title: Error Code
      type: object
      title: ErrorResponse
    ContentUploadItem:
      properties:
        source_id:
          anyOf:
            - type: string
            - type: 'null'
          title: Source Id
          description: Optional file ID (used for idempotency or re-uploads)
          example: CortexDoc1234
        title:
          anyOf:
            - type: string
            - type: 'null'
          title: Title
          description: Title or name for this content item
        content:
          type: string
          title: Content
          description: Text or markdown content to be indexed.
          examples:
            - |-
              # Introduction

              This is the document content.
          example: <content>
        is_markdown:
          type: boolean
          title: Is Markdown
          description: Whether the content is markdown formatted.
          default: false
          example: true
        tenant_metadata:
          anyOf:
            - type: string
            - type: 'null'
          title: Tenant Metadata
          description: >+
            JSON string containing tenant-level document metadata (e.g.,
            department, compliance_tag)


            Example: > "{"department":"Finance","compliance_tag":"GDPR"}"

          default: ''
        document_metadata:
          anyOf:
            - type: string
            - type: 'null'
          title: Document Metadata
          description: >+
            JSON string containing document-specific metadata (e.g., title,
            author, file_id). If file_id is not provided, the system will
            generate an ID automatically.


            Example: > "{"title":"Q1 Report.pdf","author":"Alice
            Smith","file_id":"custom_file_123"}"


          default: ''
        relations:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Relations
          description: Optional JSON string containing relations as key-value pairs
          default: false
      type: object
      required:
        - content
      title: ContentUploadItem
      description: |-
        Represents a single file/content upload.
        This model is intended to be used inside a list for batch uploads.
    SourceStatus:
      type: string
      enum:
        - queued
        - processing
        - completed
        - failed
      title: SourceStatus
  securitySchemes:
    HTTPBearer:
      type: http
      scheme: bearer

````