Vector stores
Vector stores let you upload documents, automatically chunk and embed them, and perform semantic similarity search. They are the building blocks for retrieval-augmented generation (RAG), document Q&A, knowledge bases, and semantic search applications.
How vector stores work
- Create a store — choose an embedding model, dimensions, and distance metric.
- Upload files — PDFs, text, markdown, CSV, HTML, and JSON are supported. Files are automatically parsed, split into chunks, and embedded.
- Search — query with natural language text or a raw embedding vector. Results are ranked by similarity.
Creating a vector store
import requests
response = requests.post(
"https://api.scx.ai/v1/vector-stores",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"name": "my-knowledge-base",
"embeddingModel": "E5-Mistral-7B-Instruct",
"dimensions": 4096,
"distance": "Cosine",
},
)
store = response.json()
print(store["id"], store["status"]) # pending
import requests
response = requests.post(
"https://api.scx.ai/v1/vector-stores",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"name": "my-knowledge-base",
"embeddingModel": "E5-Mistral-7B-Instruct",
"dimensions": 4096,
"distance": "Cosine",
},
)
store = response.json()
print(store["id"], store["status"]) # pending
| Parameter | Required | Description |
|---|---|---|
name | Yes | Name for the store (1-255 characters). |
embeddingModel | Yes | The embedding model to use (e.g. "E5-Mistral-7B-Instruct"). |
dimensions | Yes | Vector dimensions — must match the model's output dimensions. |
description | No | Optional description. |
distance | No | Distance metric: "Cosine" (default), "Euclid", or "Dot". |
A newly created store starts with status "pending" while the backing vector collection is provisioned. Wait for status "ready" before uploading files or searching.
Uploading files
Upload documents to a vector store using multipart form data. Each file is automatically parsed, split into chunks, and embedded in the background.
import requests
store_id = "a1b2c3d4-e5f6-7890-abcd-ef1234567890"
with open("documentation.pdf", "rb") as f:
response = requests.post(
f"https://api.scx.ai/v1/vector-stores/{store_id}/files",
headers={"Authorization": "Bearer your-scx-api-key"},
files={"file": ("documentation.pdf", f, "application/pdf")},
)
file_obj = response.json()
print(file_obj["id"], file_obj["status"]) # pending
import requests
store_id = "a1b2c3d4-e5f6-7890-abcd-ef1234567890"
with open("documentation.pdf", "rb") as f:
response = requests.post(
f"https://api.scx.ai/v1/vector-stores/{store_id}/files",
headers={"Authorization": "Bearer your-scx-api-key"},
files={"file": ("documentation.pdf", f, "application/pdf")},
)
file_obj = response.json()
print(file_obj["id"], file_obj["status"]) # pending
Supported file types
| Content Type | Extensions |
|---|---|
text/plain | .txt |
text/markdown | .md |
text/csv | .csv |
text/html | .html |
application/json | .json |
application/pdf | .pdf |
Files are processed asynchronously. After uploading, a file moves through "pending" → "processing" → "completed" (or "failed" if an error occurs).
Checking file status
List all files in a store to monitor processing progress:
response = requests.get(
f"https://api.scx.ai/v1/vector-stores/{store_id}/files",
headers={"Authorization": "Bearer your-scx-api-key"},
)
for f in response.json()["data"]:
print(f["filename"], f["status"], f["chunkCount"])
response = requests.get(
f"https://api.scx.ai/v1/vector-stores/{store_id}/files",
headers={"Authorization": "Bearer your-scx-api-key"},
)
for f in response.json()["data"]:
print(f["filename"], f["status"], f["chunkCount"])
Searching
Once files are processed, search with a natural language query. The query is automatically embedded using the store's embedding model.
response = requests.post(
f"https://api.scx.ai/v1/vector-stores/{store_id}/search",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"query": "How does authentication work?",
"limit": 5,
},
)
for result in response.json()["data"]:
print(f"[{result['score']:.2f}] {result['payload']['filename']}")
print(f" {result['payload']['text'][:200]}")
response = requests.post(
f"https://api.scx.ai/v1/vector-stores/{store_id}/search",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"query": "How does authentication work?",
"limit": 5,
},
)
for result in response.json()["data"]:
print(f"[{result['score']:.2f}] {result['payload']['filename']}")
print(f" {result['payload']['text'][:200]}")
Search parameters
| Parameter | Required | Description |
|---|---|---|
query | No* | Text query — automatically embedded. *At least one of query or vector is required. |
vector | No* | Raw embedding vector. *At least one of query or vector is required. |
limit | No | Maximum number of results (1-100, default 10). |
scoreThreshold | No | Minimum similarity score to include in results. |
Example response
{
"data": [
{
"id": "abc123",
"score": 0.92,
"payload": {
"text": "Authentication is handled via API keys passed in the Authorization header...",
"filename": "auth-docs.pdf"
}
},
{
"id": "def456",
"score": 0.87,
"payload": {
"text": "OAuth 2.0 providers can be configured for single sign-on...",
"filename": "sso-guide.pdf"
}
}
]
}
{
"data": [
{
"id": "abc123",
"score": 0.92,
"payload": {
"text": "Authentication is handled via API keys passed in the Authorization header...",
"filename": "auth-docs.pdf"
}
},
{
"id": "def456",
"score": 0.87,
"payload": {
"text": "OAuth 2.0 providers can be configured for single sign-on...",
"filename": "sso-guide.pdf"
}
}
]
}
Searching with a raw vector
If you have a pre-computed embedding, pass it directly instead of a text query:
response = requests.post(
f"https://api.scx.ai/v1/vector-stores/{store_id}/search",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"vector": [0.0123, -0.0456, 0.0789, ...], # your embedding
"limit": 5,
},
)
response = requests.post(
f"https://api.scx.ai/v1/vector-stores/{store_id}/search",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"vector": [0.0123, -0.0456, 0.0789, ...], # your embedding
"limit": 5,
},
)
Managing stores and files
Listing stores
response = requests.get(
"https://api.scx.ai/v1/vector-stores",
headers={"Authorization": "Bearer your-scx-api-key"},
)
for store in response.json()["data"]:
print(store["name"], store["status"], store["pointCount"], "vectors")
response = requests.get(
"https://api.scx.ai/v1/vector-stores",
headers={"Authorization": "Bearer your-scx-api-key"},
)
for store in response.json()["data"]:
print(store["name"], store["status"], store["pointCount"], "vectors")
Deleting a file
response = requests.delete(
f"https://api.scx.ai/v1/vector-stores/{store_id}/files/{file_id}",
headers={"Authorization": "Bearer your-scx-api-key"},
)
print(response.json()) # {"deleted": true}
response = requests.delete(
f"https://api.scx.ai/v1/vector-stores/{store_id}/files/{file_id}",
headers={"Authorization": "Bearer your-scx-api-key"},
)
print(response.json()) # {"deleted": true}
Deleting a store
response = requests.delete(
f"https://api.scx.ai/v1/vector-stores/{store_id}",
headers={"Authorization": "Bearer your-scx-api-key"},
)
print(response.json()) # {"deleted": true}
response = requests.delete(
f"https://api.scx.ai/v1/vector-stores/{store_id}",
headers={"Authorization": "Bearer your-scx-api-key"},
)
print(response.json()) # {"deleted": true}
Deleting a store permanently removes all associated files and vectors. This action cannot be undone.
Distance metrics
Choose a distance metric when creating a store. It cannot be changed after creation.
| Metric | Best for |
|---|---|
Cosine | Text embeddings (default). Measures angular distance, normalized to [-1, 1]. |
Euclid | Dense spatial data. Measures straight-line distance between points. |
Dot | Scenarios where vector magnitude carries meaning. |
For most text-based use cases, Cosine is the best default.
End-to-end example
A complete example that creates a store, uploads a file, waits for processing, and searches:
import requests
import time
BASE_URL = "https://api.scx.ai/v1"
HEADERS = {"Authorization": "Bearer your-scx-api-key"}
# 1. Create the store
store = requests.post(
f"{BASE_URL}/vector-stores",
headers=HEADERS,
json={
"name": "product-docs",
"embeddingModel": "E5-Mistral-7B-Instruct",
"dimensions": 4096,
},
).json()
print(f"Store created: {store['id']}")
# 2. Wait for the store to be ready
while store["status"] == "pending":
time.sleep(2)
store = requests.get(
f"{BASE_URL}/vector-stores/{store['id']}", headers=HEADERS
).json()
print(f"Store status: {store['status']}")
# 3. Upload a file
with open("docs.pdf", "rb") as f:
file_obj = requests.post(
f"{BASE_URL}/vector-stores/{store['id']}/files",
headers=HEADERS,
files={"file": ("docs.pdf", f, "application/pdf")},
).json()
print(f"File uploaded: {file_obj['id']}")
# 4. Wait for processing
while file_obj["status"] in ("pending", "processing"):
time.sleep(5)
files = requests.get(
f"{BASE_URL}/vector-stores/{store['id']}/files", headers=HEADERS
).json()
file_obj = next(f for f in files["data"] if f["id"] == file_obj["id"])
print(f"File status: {file_obj['status']} ({file_obj['chunkCount']} chunks)")
# 5. Search
results = requests.post(
f"{BASE_URL}/vector-stores/{store['id']}/search",
headers=HEADERS,
json={"query": "How do I get started?", "limit": 3},
).json()
for r in results["data"]:
print(f"[{r['score']:.2f}] {r['payload']['text'][:150]}")
import requests
import time
BASE_URL = "https://api.scx.ai/v1"
HEADERS = {"Authorization": "Bearer your-scx-api-key"}
# 1. Create the store
store = requests.post(
f"{BASE_URL}/vector-stores",
headers=HEADERS,
json={
"name": "product-docs",
"embeddingModel": "E5-Mistral-7B-Instruct",
"dimensions": 4096,
},
).json()
print(f"Store created: {store['id']}")
# 2. Wait for the store to be ready
while store["status"] == "pending":
time.sleep(2)
store = requests.get(
f"{BASE_URL}/vector-stores/{store['id']}", headers=HEADERS
).json()
print(f"Store status: {store['status']}")
# 3. Upload a file
with open("docs.pdf", "rb") as f:
file_obj = requests.post(
f"{BASE_URL}/vector-stores/{store['id']}/files",
headers=HEADERS,
files={"file": ("docs.pdf", f, "application/pdf")},
).json()
print(f"File uploaded: {file_obj['id']}")
# 4. Wait for processing
while file_obj["status"] in ("pending", "processing"):
time.sleep(5)
files = requests.get(
f"{BASE_URL}/vector-stores/{store['id']}/files", headers=HEADERS
).json()
file_obj = next(f for f in files["data"] if f["id"] == file_obj["id"])
print(f"File status: {file_obj['status']} ({file_obj['chunkCount']} chunks)")
# 5. Search
results = requests.post(
f"{BASE_URL}/vector-stores/{store['id']}/search",
headers=HEADERS,
json={"query": "How do I get started?", "limit": 3},
).json()
for r in results["data"]:
print(f"[{r['score']:.2f}] {r['payload']['text'][:150]}")
Using vector stores with chat (RAG)
A common pattern is retrieval-augmented generation: search the vector store for relevant context, then pass it to a chat completion. This grounds the model's response in your documents.
from openai import OpenAI
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
# Search the vector store for context
import requests
search_results = requests.post(
f"https://api.scx.ai/v1/vector-stores/{store_id}/search",
headers={"Authorization": "Bearer your-scx-api-key"},
json={"query": "How do I reset my password?", "limit": 3},
).json()
# Build context from search results
context = "\n\n".join(r["payload"]["text"] for r in search_results["data"])
# Use context in a chat completion
response = client.chat.completions.create(
model="Meta-Llama-3.3-70B-Instruct",
messages=[
{
"role": "system",
"content": f"Answer the user's question using the following context:\n\n{context}",
},
{"role": "user", "content": "How do I reset my password?"},
],
)
print(response.choices[0].message.content)
from openai import OpenAI
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
# Search the vector store for context
import requests
search_results = requests.post(
f"https://api.scx.ai/v1/vector-stores/{store_id}/search",
headers={"Authorization": "Bearer your-scx-api-key"},
json={"query": "How do I reset my password?", "limit": 3},
).json()
# Build context from search results
context = "\n\n".join(r["payload"]["text"] for r in search_results["data"])
# Use context in a chat completion
response = client.chat.completions.create(
model="Meta-Llama-3.3-70B-Instruct",
messages=[
{
"role": "system",
"content": f"Answer the user's question using the following context:\n\n{context}",
},
{"role": "user", "content": "How do I reset my password?"},
],
)
print(response.choices[0].message.content)
Store and file statuses
Store statuses
| Status | Description |
|---|---|
pending | The vector collection is being provisioned. |
ready | The store is ready to accept file uploads and search queries. |
error | Provisioning failed. Try creating a new store. |
File statuses
| Status | Description |
|---|---|
pending | The file is queued for processing. |
processing | The file is being parsed, chunked, and embedded. |
completed | The file has been successfully indexed and is searchable. |
failed | Processing failed. Check the error field for details. |
For detailed endpoint documentation, see the Vector Stores API reference.