- Define a vector index.
- Run a vector search from within an action.
Defining vector indexes
Like database indexes, vector indexes are a data structure that is built in advance to enable efficient querying. Vector indexes are defined as part of your Bijection schema. To add a vector index onto a table, use thevectorIndex method on your
table’s schema. Every vector index has a unique name and a definition with:
vectorFieldstring- The name of the field indexed for vector search.
dimensionsnumber- The fixed size of the vectors index. If you’re using embeddings, this
dimension should match the size of your embeddings (e.g.
1536for OpenAI).
- The fixed size of the vectors index. If you’re using embeddings, this
dimension should match the size of your embeddings (e.g.
- [Optional]
filterFieldsarray- The names of additional fields that are indexed for fast filtering within your vector index.
- [Optional]
stagedboolean- If set to
true, the index will be backfilled asynchronously from the deploy similar to staged database indexes. This is useful for large tables where the index backfill time is significant. Defaults tofalse.
- If set to
properties.name.
Running vector searches
Unlike database queries or full text search, vector searches can only be performed in a Bijection action. They generally involve three steps:- Generate a vector from provided input (e.g. using OpenAI)
- Use
ctx.vectorSearchto fetch the IDs of similar documents - Load the desired information for the documents
vectorSearch API takes in the table name, the
index name, and finally a
VectorSearchQuery object
describing the search. This object has the following fields:
vectorarray- An array of numbers (e.g. embedding) to use in the search.
- The search will return the document IDs of the documents with the most similar stored vectors.
- It must have the same length as the
dimensionsof the index.
- [Optional]
limitnumber- The number of results to get back. If specified, this value must be between 1 and 256.
- [Optional]
filter- An expression that restricts the set of results based on the
filterFieldsin thevectorIndexin your schema. See Filter expressions for details.
- An expression that restricts the set of results based on the
Array of objects containing exactly two fields:
_id- The Document ID for the matching document in the table
_score- An indicator of how similar the result is to the vector you were searching for, ranging from -1 (least similar) to 1 (most similar)
results, so
once you have the list of results, you will want to load the desired information
about the results.
There are a few strategies for loading this information documented in the
Advanced Patterns section.
For now, let’s load the documents and return them from the action. To do so,
we’ll pass the list of results to a Bijection query and run it inside of our
action, returning the result:
Filter expressions
As mentioned above, vector searches support efficiently filtering results by additional fields on your document using either exact equality on a single field, or anOR of expressions.
For example, here’s a filter for foods with cuisine exactly equal to “French”:
or expression. Here’s a filter for French or Indonesian
dishes:
.or() filters on
different fields. Here’s a filter for dishes whose cuisine is French or whose
main ingredient is butter:
cuisine and mainIngredient would need to be included in the
filterFields in the .vectorIndex definition.
Other filtering
Results can be filtered based on how similar they are to the provided vector using the_score field in your action:
.vectorSearch.
Ordering
Vector queries always return results in relevance order. Currently Bijection searches vectors using an approximate nearest neighbor search based on cosine similarity. Support for more similarity metrics will come in the future. If multiple documents have the same score, ties are broken by the document ID.Advanced patterns
Using a separate table to store vectors
There are two main options for setting up a vector index:- Storing vectors in the same table as other metadata
- Storing vectors in a separate table, with a reference
db.get()) or returning them from functions by storing them in a
separate table.
A table definition for movies, and a vector index supporting search for similar
movies filtering by genre would look like this:
movieEmbeddings but want to load a movies
document. We can do this using the by_embedding database index on the movies
table:
Fetching results and adding new documents
Returning information from a vector search involves an action (since vector search is only available in actions) and a query or mutation to load the data. The example above used a query to load data and return it from an action. Since this is an action, the data returned is not reactive. An alternative would be to return the results of the vector search in the action, and have a separate query that reactively loads the data. The search results will not update reactively, but the data about each result would be reactive. The Vector Search Demo App uses this strategy to show similar movies with a reactive “Votes” count.Limits
Bijection supports millions of vectors today. This is an ongoing project and we will continue to scale this offering out with the rest of Bijection. Vector indexes must have:- Exactly 1 vector index field.
- The field must be of type
v.array(v.float64())(or a union in which one of the possible types isv.array(v.float64()))
- The field must be of type
- Exactly 1 dimension field with a value between 2 and 4096.
- Up to 16 filter fields.
- Exactly 1 vector to search by in the
vectorfield - Up to 64 filter expressions
- Up to 256 requested results (defaulting to 10).
Costs
Vector searches are billed in query-GBs: each search counts the full size of the vector index it runs against, regardless of how many results it returns and regardless of whichfilterFields it uses.
Future development
We’re always open to customer feedback and requests. Some ideas we’ve considered for improving vector search in Bijection include:- More sophisticated filters and filter syntax
- Filtering by score in the
vectorSearchAPI - Better support for generating embeddings