Vector stores

Vector stores let you upload documents, automatically chunk and embed them, and perform semantic similarity search. They are the building blocks for retrieval-augmented generation (RAG), document Q&A, knowledge bases, and semantic search applications.

How vector stores work

  1. Create a store — choose an embedding model, dimensions, and distance metric.
  2. Upload files — PDFs, text, markdown, CSV, HTML, and JSON are supported. Files are automatically parsed, split into chunks, and embedded.
  3. Search — query with natural language text or a raw embedding vector. Results are ranked by similarity.

Creating a vector store

import requests

response = requests.post(
    "https://api.scx.ai/v1/vector-stores",
    headers={"Authorization": "Bearer your-scx-api-key"},
    json={
        "name": "my-knowledge-base",
        "embeddingModel": "E5-Mistral-7B-Instruct",
        "dimensions": 4096,
        "distance": "Cosine",
    },
)

store = response.json()
print(store["id"], store["status"])  # pending
ParameterRequiredDescription
nameYesName for the store (1-255 characters).
embeddingModelYesThe embedding model to use (e.g. "E5-Mistral-7B-Instruct").
dimensionsYesVector dimensions — must match the model's output dimensions.
descriptionNoOptional description.
distanceNoDistance metric: "Cosine" (default), "Euclid", or "Dot".

Uploading files

Upload documents to a vector store using multipart form data. Each file is automatically parsed, split into chunks, and embedded in the background.

import requests

store_id = "a1b2c3d4-e5f6-7890-abcd-ef1234567890"

with open("documentation.pdf", "rb") as f:
    response = requests.post(
        f"https://api.scx.ai/v1/vector-stores/{store_id}/files",
        headers={"Authorization": "Bearer your-scx-api-key"},
        files={"file": ("documentation.pdf", f, "application/pdf")},
    )

file_obj = response.json()
print(file_obj["id"], file_obj["status"])  # pending

Supported file types

Content TypeExtensions
text/plain.txt
text/markdown.md
text/csv.csv
text/html.html
application/json.json
application/pdf.pdf

Files are processed asynchronously. After uploading, a file moves through "pending" → "processing" → "completed" (or "failed" if an error occurs).

Checking file status

List all files in a store to monitor processing progress:

response = requests.get(
    f"https://api.scx.ai/v1/vector-stores/{store_id}/files",
    headers={"Authorization": "Bearer your-scx-api-key"},
)

for f in response.json()["data"]:
    print(f["filename"], f["status"], f["chunkCount"])

Searching

Once files are processed, search with a natural language query. The query is automatically embedded using the store's embedding model.

response = requests.post(
    f"https://api.scx.ai/v1/vector-stores/{store_id}/search",
    headers={"Authorization": "Bearer your-scx-api-key"},
    json={
        "query": "How does authentication work?",
        "limit": 5,
    },
)

for result in response.json()["data"]:
    print(f"[{result['score']:.2f}] {result['payload']['filename']}")
    print(f"  {result['payload']['text'][:200]}")

Search parameters

ParameterRequiredDescription
queryNo*Text query — automatically embedded. *At least one of query or vector is required.
vectorNo*Raw embedding vector. *At least one of query or vector is required.
limitNoMaximum number of results (1-100, default 10).
scoreThresholdNoMinimum similarity score to include in results.

Example response

json
{
  "data": [
    {
      "id": "abc123",
      "score": 0.92,
      "payload": {
        "text": "Authentication is handled via API keys passed in the Authorization header...",
        "filename": "auth-docs.pdf"
      }
    },
    {
      "id": "def456",
      "score": 0.87,
      "payload": {
        "text": "OAuth 2.0 providers can be configured for single sign-on...",
        "filename": "sso-guide.pdf"
      }
    }
  ]
}

Searching with a raw vector

If you have a pre-computed embedding, pass it directly instead of a text query:

response = requests.post(
    f"https://api.scx.ai/v1/vector-stores/{store_id}/search",
    headers={"Authorization": "Bearer your-scx-api-key"},
    json={
        "vector": [0.0123, -0.0456, 0.0789, ...],  # your embedding
        "limit": 5,
    },
)

Managing stores and files

Listing stores

response = requests.get(
    "https://api.scx.ai/v1/vector-stores",
    headers={"Authorization": "Bearer your-scx-api-key"},
)

for store in response.json()["data"]:
    print(store["name"], store["status"], store["pointCount"], "vectors")

Deleting a file

response = requests.delete(
    f"https://api.scx.ai/v1/vector-stores/{store_id}/files/{file_id}",
    headers={"Authorization": "Bearer your-scx-api-key"},
)

print(response.json())  # {"deleted": true}

Deleting a store

response = requests.delete(
    f"https://api.scx.ai/v1/vector-stores/{store_id}",
    headers={"Authorization": "Bearer your-scx-api-key"},
)

print(response.json())  # {"deleted": true}

Distance metrics

Choose a distance metric when creating a store. It cannot be changed after creation.

MetricBest for
CosineText embeddings (default). Measures angular distance, normalized to [-1, 1].
EuclidDense spatial data. Measures straight-line distance between points.
DotScenarios where vector magnitude carries meaning.

For most text-based use cases, Cosine is the best default.

End-to-end example

A complete example that creates a store, uploads a file, waits for processing, and searches:

import requests
import time

BASE_URL = "https://api.scx.ai/v1"
HEADERS = {"Authorization": "Bearer your-scx-api-key"}

# 1. Create the store
store = requests.post(
    f"{BASE_URL}/vector-stores",
    headers=HEADERS,
    json={
        "name": "product-docs",
        "embeddingModel": "E5-Mistral-7B-Instruct",
        "dimensions": 4096,
    },
).json()
print(f"Store created: {store['id']}")

# 2. Wait for the store to be ready
while store["status"] == "pending":
    time.sleep(2)
    store = requests.get(
        f"{BASE_URL}/vector-stores/{store['id']}", headers=HEADERS
    ).json()
print(f"Store status: {store['status']}")

# 3. Upload a file
with open("docs.pdf", "rb") as f:
    file_obj = requests.post(
        f"{BASE_URL}/vector-stores/{store['id']}/files",
        headers=HEADERS,
        files={"file": ("docs.pdf", f, "application/pdf")},
    ).json()
print(f"File uploaded: {file_obj['id']}")

# 4. Wait for processing
while file_obj["status"] in ("pending", "processing"):
    time.sleep(5)
    files = requests.get(
        f"{BASE_URL}/vector-stores/{store['id']}/files", headers=HEADERS
    ).json()
    file_obj = next(f for f in files["data"] if f["id"] == file_obj["id"])
print(f"File status: {file_obj['status']} ({file_obj['chunkCount']} chunks)")

# 5. Search
results = requests.post(
    f"{BASE_URL}/vector-stores/{store['id']}/search",
    headers=HEADERS,
    json={"query": "How do I get started?", "limit": 3},
).json()

for r in results["data"]:
    print(f"[{r['score']:.2f}] {r['payload']['text'][:150]}")

Using vector stores with chat (RAG)

A common pattern is retrieval-augmented generation: search the vector store for relevant context, then pass it to a chat completion. This grounds the model's response in your documents.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.scx.ai/v1",
    api_key="your-scx-api-key",
)

# Search the vector store for context
import requests

search_results = requests.post(
    f"https://api.scx.ai/v1/vector-stores/{store_id}/search",
    headers={"Authorization": "Bearer your-scx-api-key"},
    json={"query": "How do I reset my password?", "limit": 3},
).json()

# Build context from search results
context = "\n\n".join(r["payload"]["text"] for r in search_results["data"])

# Use context in a chat completion
response = client.chat.completions.create(
    model="Meta-Llama-3.3-70B-Instruct",
    messages=[
        {
            "role": "system",
            "content": f"Answer the user's question using the following context:\n\n{context}",
        },
        {"role": "user", "content": "How do I reset my password?"},
    ],
)

print(response.choices[0].message.content)

Store and file statuses

Store statuses

StatusDescription
pendingThe vector collection is being provisioned.
readyThe store is ready to accept file uploads and search queries.
errorProvisioning failed. Try creating a new store.

File statuses

StatusDescription
pendingThe file is queued for processing.
processingThe file is being parsed, chunked, and embedded.
completedThe file has been successfully indexed and is searchable.
failedProcessing failed. Check the error field for details.

For detailed endpoint documentation, see the Vector Stores API reference.