Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
39 changes: 38 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,40 @@ Repository: https://github.com/asynq-io/sqlargon

---

## Features

- **Repository pattern** — one object wraps async sessions, core queries and ORM models;
sessions are context-local and resolved at call time, so nothing gets passed around
- **High-level CRUD** — `create`, `get`, `get_or_create`, `create_or_update`, `all`,
`list`, `count`, `update_one`, `update_many`, `delete_one`, `delete_many` and `remove`
out of the box
- **Bulk operations** — `bulk_create`, `bulk_create_or_update` and `bulk_update` with
per-repository conflict handling
- **Query builder** — fluent, dialect-aware statements for upserts, `RETURNING`, advisory
locks and streaming, with terminal helpers that cast results to `.scalars()`, `.one()`,
`.mappings()`, ...
- **Multi-dialect** — PostgreSQL, SQLite, MySQL and MariaDB, with capability-gated SQL
generation per backend
- **Transactions** — `@atomic` and database-scoped `atomic()` blocks, plus named advisory
locks
- **Unit of work** — repositories declared as annotations on a unit of work share one
session and one transaction
- **Database routing** — clusters with read replicas, shards and vertical partitioning;
`using()`, `read_only` and per-request `use_context`
- **Pagination** — page-number, offset/limit and keyset cursor strategies
- **Outbox** — transactional outbox with a background relay and eventiq integration
- **Cron** — database-backed scheduler with namespaces and safe multi-instance claiming
- **Column types and mixins** — UUID (v4/v7), timestamp, orjson JSON and pydantic-validated
columns; mixins for UUID keys, created/updated timestamps and soft delete
- **Soft delete** — tombstone-based deletes via `SoftDeleteRepository`
- **Versioned models** — optimistic concurrency with UUID or PostgreSQL `xmin` versions
- **Auditable models** — append-only versioned history with point-in-time reads and restore
- **Vector search** — embeddings with cosine, L2, dot and L1 similarity, full-text and
hybrid reciprocal-rank-fusion search on PostgreSQL and SQLite
- **FastAPI-ready** — repositories and units of work work directly as dependencies
- **Alembic migrations** — async-first migration setup
- **OpenTelemetry** — optional SQLAlchemy instrumentation

## About

This library provides glue code to use sqlalchemy async sessions, core queries and orm models
Expand All @@ -37,6 +71,7 @@ from one object which provides somewhat of repository pattern. This solution has
- engines and routing policy are separate, so the same repository runs against one database,
a primary with read replicas, or a set of shards


## Installation

```shell
Expand Down Expand Up @@ -331,7 +366,9 @@ ordering.

`sqlargon.types` provides dialect-aware column types: `GUID` with `GenerateUUID` /
`GenerateUUIDV7` server defaults, `Timestamp` with a `now()` server default and `JSON`
(orjson-serialized). `sqlargon.types.pydantic` adds `Pydantic` and `ValidatedType` for
(orjson-serialized), whose comparator carries portable JSON operators — containment and
key tests, plus server-side mutation (`set_key`, `update`, `remove_key`) that rewrites a
document in the `UPDATE` itself. `sqlargon.types.pydantic` adds `Pydantic` and `ValidatedType` for
pydantic-validated columns. `sqlargon.mixins` bundles them into `UUIDModelMixin`,
`UUIDV7ModelMixin`, `CreatedUpdatedMixin` and `SoftDeleteMixin`.

Expand Down
57 changes: 40 additions & 17 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,31 +13,50 @@
*SQLAlchemy repository pattern and utilities*

---
Version: 1.0.0b1

Docs: [https://asynq-io.github.io/sqlargon/](https://asynq-io.github.io/sqlargon/)

Repository: [https://github.com/asynq-io/sqlargon](https://github.com/asynq-io/sqlargon)

---

## About
## Features

- **Repository pattern** — one object wraps async sessions, core queries and ORM models;
sessions are context-local and resolved at call time, so nothing gets passed around
- **High-level CRUD** — `create`, `get`, `get_or_create`, `create_or_update`, `all`,
`list`, `count`, `update_one`, `update_many`, `delete_one`, `delete_many` and `remove`
out of the box
- **Bulk operations** — `bulk_create`, `bulk_create_or_update` and `bulk_update` with
per-repository conflict handling
- **Query builder** — fluent, dialect-aware statements for upserts, `RETURNING`, advisory
locks and streaming, with terminal helpers that cast results to `.scalars()`, `.one()`,
`.mappings()`, ...
- **Multi-dialect** — PostgreSQL, SQLite, MySQL and MariaDB, with capability-gated SQL
generation per backend
- **Transactions** — `@atomic` and database-scoped `atomic()` blocks, plus named advisory
locks
- **Unit of work** — repositories declared as annotations on a unit of work share one
session and one transaction
- **Database routing** — [clusters](routing.md) with read replicas, shards and vertical
partitioning; `using()`, `read_only` and per-request `use_context`
- **Pagination** — [page-number, offset/limit and keyset cursor](pagination.md) strategies
- **Outbox** — [transactional outbox](outbox.md) with a background relay and eventiq
integration
- **Cron** — [database-backed scheduler](cron.md) with namespaces and safe multi-instance
claiming
- **Column types and mixins** — UUID (v4/v7), timestamp, orjson JSON and pydantic-validated
columns; mixins for UUID keys, created/updated timestamps and soft delete
- **Soft delete** — tombstone-based deletes via `SoftDeleteRepository`
- **Versioned models** — optimistic concurrency with UUID or PostgreSQL `xmin` versions
- **Auditable models** — [append-only versioned history](auditable.md) with point-in-time
reads and restore
- **Vector search** — [embeddings with similarity, full-text and hybrid
reciprocal-rank-fusion search](vectors.md) on PostgreSQL and SQLite
- **FastAPI-ready** — repositories and units of work work directly as dependencies
- **Alembic migrations** — async-first [migration setup](migrations.md)
- **OpenTelemetry** — optional SQLAlchemy instrumentation

SQLArgon provides glue code to use SQLAlchemy async sessions, core queries and ORM models
from one object which provides somewhat of a repository pattern. This solution has a few
advantages:

- no need to pass a `session` object to every function/method — sessions are context-local
and resolved by the repository itself
- write data access queries in one place
- no need to import `insert`, `update`, `delete`, `select` from SQLAlchemy over and over again
- implicit cast of results to `.scalars().all()`, `.one()`, `.mappings()`, ...
- a dialect-aware query builder for upserts, `RETURNING` and advisory locks
- your view model (e.g. FastAPI routes) does not need to know about the underlying storage —
the repository class can be replaced at any moment with any object providing a similar
interface
- engines and routing policy are separate, so the same repository runs against one database,
a primary with read replicas, or a set of shards

## Installation

Expand Down Expand Up @@ -113,6 +132,10 @@ or from `DATABASE_*` environment variables.
- **[Usage](usage.md)** — models, CRUD, query building, transactions and units of work.
- **[Database Routing](routing.md)** — replicas, shards, routers and FastAPI wiring.
- **[Pagination](pagination.md)** — page-number, offset/limit and cursor strategies.
- **[Cron](cron.md)** — database-backed scheduling with namespaces and multi-instance safety.
- **[Outbox](outbox.md)** — the transactional outbox pattern and its relay.
- **[Vector Search](vectors.md)** — embeddings, similarity and hybrid search.
- **[Auditable Models](auditable.md)** — append-only versioned history.
- **[Examples](examples.md)** — end-to-end recipes: a FastAPI service, batch workers,
multi-tenant sharding, testing.
- **Reference** — [types and mixins](reference/types.md), [dialects](reference/dialects.md),
Expand Down
37 changes: 37 additions & 0 deletions docs/outbox.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,43 @@ The topic, the event type of each recorded operation and the payload columns
are derived from the config, and readable for inspection as `repository.topic`,
`repository.event_types` and `repository.payload_columns`.

### Topic templating

A `topic` may carry `{placeholders}` naming attributes of the written row, for
a topic that identifies its subject rather than the table that holds it — say
one topic per tenant or per entity. Each placeholder is filled from the row the
event was written from, at write time, using `str.format`:

```python
class UserRepository(OutboxRepository[User]):
outbox = OutboxConfig(
topic="events.organizations.{organization_id}.deleted",
exclude={"password"},
)


await UserRepository().create(
name="John", password=hashed, organization_id=21
)
# -> topic "events.organizations.21.deleted"
```

A topic without placeholders is used verbatim, so nothing changes for topics
that do not use them. A placeholder names a column — or any other attribute of
the row — and its value is stringified as it is: a `UUID` becomes its
canonical string form. `{id}` reads the row's `id`, `{organization_id}` its
`organization_id`, and so on.

The value is read the same way the payload and the extra attributes are: from
the row the write produced (before it, for a delete). It is read **when the
write happens**, not when the relay publishes the event, so a templated topic
always reflects the state at write time. A repository serves one event per
written row, so a bulk write of rows from different organizations lands on
their own topics.

`format_topic(topic, row)` does the substitution on its own and is exported
from `sqlargon.outbox`.

### Extra attributes

Some values belong *next to* the payload rather than inside it — a `tenant_id`
Expand Down
2 changes: 2 additions & 0 deletions docs/reference/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -126,6 +126,8 @@ Generated from the source. See [Usage](../usage.md) for a narrative introduction

::: sqlargon.outbox.Operation

::: sqlargon.outbox.format_topic

::: sqlargon.integrations.eventiq.to_cloud_event

::: sqlargon.integrations.eventiq.eventiq_publisher
Expand Down
19 changes: 13 additions & 6 deletions docs/reference/dialects.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,12 +11,12 @@ is isolated.
Builders declare what they support as an `Option` flag, checked with
`db.query_builder.supports(...)`:

| Dialect | `RETURNING` | `CONFLICTS` | `LOCKS` |
| --- | --- | --- | --- |
| `postgresql` | ✅ | ✅ | ✅ `pg_advisory_lock` |
| `sqlite` | ✅ (SQLite ≥ 3.35) | ✅ | ❌ |
| `mysql` | ❌ | ✅ | ✅ `GET_LOCK` |
| anything else | ❌ | ❌ | ❌ |
| Dialect | `RETURNING` | `CONFLICTS` | `LOCKS` | `VECTORS` | `FULL_TEXT` |
| --- | --- | --- | --- | --- | --- |
| `postgresql` | ✅ | ✅ | ✅ `pg_advisory_lock` | ✅ pgvector | ✅ `ts_rank` |
| `sqlite` | ✅ (SQLite ≥ 3.35) | ✅ | ❌ | ✅ sqlite-vector | ❌ |
| `mysql` | ❌ | ✅ | ✅ `GET_LOCK` | ❌ | ❌ |
| anything else | ❌ | ❌ | ❌ | ❌ | ❌ |

```python
from sqlargon.query_builder import Option
Expand Down Expand Up @@ -106,6 +106,13 @@ Methods: `select`, `insert`, `update`, `delete`, `filter`, `count`, `page`, `loc
and `get_lock_pair`. `insert`, `update` and `delete` take `return_results=True` to append a
`RETURNING` clause for the whole table.

The search hooks — `vector_search`, `vector_distance`, `vector_init`, `text_search` and
`rrf_search` — are the same idea for [vector search](../vectors.md): the base class refuses
them with `UnsupportedDialectError`, and the two backends that can search express it in
shapes with nothing in common. PostgreSQL orders by a pgvector operator; SQLite joins the
table valued scan sqlite-vector exposes, because it has no scalar distance function at all.
Keeping both behind one hook is what lets `VectorRepository.search()` be portable.

## Adding a dialect

Subclass `QueryBuilder`, declare the supported options and override what differs — the base
Expand Down
61 changes: 56 additions & 5 deletions docs/reference/types.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,18 +91,69 @@ await repo.list(Document.meta.json_value("owner") == "john")
| `has_any_key([...])` | `?\|` | `JSON_CONTAINS_PATH(..., 'one', ...)` | `EXISTS` over `json_each` |
| `has_all_keys([...])` | `?&` | `JSON_CONTAINS_PATH(..., 'all', ...)` | `json_each` self-join |
| `json_value(key)` | `->>` | `JSON_EXTRACT` | `JSON_EXTRACT` |
| `get(key)` | `->` | `JSON_EXTRACT` | `JSON_EXTRACT` |
| `has_key(key)` | `?` | `JSON_CONTAINS_PATH(..., 'one', ...)` | `JSON_TYPE(...) IS NOT NULL` |
| `array_length()` | `JSONB_ARRAY_LENGTH` | `JSON_LENGTH` | `JSON_ARRAY_LENGTH` |
| `keys()` | `JSONB_OBJECT_KEYS` + `JSONB_AGG` | `JSON_KEYS` | `JSON_GROUP_ARRAY` over `json_each` |

!!! warning "`has_any_key` / `has_all_keys` are portable over arrays, not objects"

Use them to test membership in a JSON **array** — that is the one meaning all three
dialects agree on. Against a JSON **object** they diverge: PostgreSQL and MySQL test the
object's *keys*, while the SQLite fallback tests the *values* produced by `json_each`.
To query a key of an object portably, use `json_value(key)` instead.
To test a key of an object portably, use `has_key(key)`, which addresses object keys on
every dialect.

The underlying function elements — `json_contains`, `json_has_any_key`, `json_has_all_keys`
and `json_value` — are importable from `sqlargon.types.json` for use outside a `JSON`
column. `has_any_key` and `has_all_keys` require string keys and raise `ValueError`
otherwise.
### Mutating a document server-side

The mutation operators rewrite a document in the `UPDATE` itself, so a single key can be
changed without reading the row into Python and writing it back — no lost update, one
round trip:

```python
await repo.update({Document.meta: Document.meta.set_key("owner", "john")}).execute()
await repo.update({Document.meta: Document.meta.update({"owner": "john", "hits": 0})}).execute()
await repo.update({Document.meta: Document.meta.remove_key("owner")}).execute()
```

| Operator | PostgreSQL | MySQL | SQLite |
| --- | --- | --- | --- |
| `set_key(key, value)` | `\|\|` | `JSON_SET` | `JSON_SET` |
| `update({...})` | `\|\|` | `JSON_SET` | `JSON_SET` |
| `remove_key(*keys)` | `-` over `text[]` | `JSON_REMOVE` | `JSON_REMOVE` |
| `insert_key(key, value)` | `\|\|`, patch on the left | `JSON_INSERT` | `JSON_INSERT` |
| `replace_key(key, value)` | `JSONB_SET(..., false)` | `JSON_REPLACE` | `JSON_REPLACE` |
| `array_append(value)` | `\|\|` + `JSONB_BUILD_ARRAY` | `JSON_ARRAY_APPEND` | `JSON_INSERT(..., '$[#]', ...)` |

`insert_key` only writes a key that is **absent**; `replace_key` only one already
**present**. Every mutation returns a JSON expression, so they nest:

```python
Document.meta.update({"c": 3}).remove_key("a")
```

!!! warning "What the mutation operators do not smooth over"

- **`NULL` in, `NULL` out.** `JSONB_SET` and `JSON_SET` both return `NULL` for a `NULL`
document, and these operators match that rather than coalescing to `{}`. Give the
column a `server_default` of `'{}'` if you need a document to always be there.
- **Objects only.** The `JSON_SET` family addresses `$."key"`, so `set_key`,
`update`, `remove_key`, `insert_key` and `replace_key` assume the document is an
object. Use `array_append` for arrays.
- **Top-level keys only.** There are no nested paths or array indices; a key is always
one level down.
- **`update` is a shallow merge.** A top-level key is replaced wholesale, not merged
into recursively — the semantics of PostgreSQL's `||`. Deep merge-patch
(`JSON_MERGE_PATCH`, `json_patch`) is deliberately absent: PostgreSQL has no builtin
for it.
- **`array_length` is portable over arrays only.** Given an object PostgreSQL raises,
SQLite answers 0 and MySQL answers 1.

The underlying function elements — `json_contains`, `json_has_any_key`, `json_has_all_keys`,
`json_value`, `json_get`, `json_has_key`, `json_array_length`, `json_keys`, `json_update`,
`json_set_key`, `json_remove_key`, `json_insert_key`, `json_replace_key` and
`json_array_append` — are importable from `sqlargon.types.json` for use outside a `JSON`
column. The key operators require string keys and raise `ValueError` otherwise.

## Pydantic-validated columns

Expand Down
Loading
Loading