# Refget Python API Documentation

### FASTA Processing

<a id="refget.utils.fasta_to_seqcol_dict"></a>

### `fasta_to_seqcol_dict`

```python
def fasta_to_seqcol_dict(fasta_file_path: Union[str, Path]) -> dict
```

Convert a FASTA file into a Sequence Collection dict.

**Parameters:**

- `fasta_file_path` (*Union[str, Path]*): Path to the FASTA file

**Returns:**

- *dict*: A canonical sequence collection dictionary

**Raises:**

- *ImportError*: If gtars is not installed (required for FASTA processing)

<a id="refget.utils.compare_seqcols"></a>

### `compare_seqcols`

```python
def compare_seqcols(A: SeqColDict, B: SeqColDict) -> dict
```

Workhorse comparison function

**Parameters:**

- `A` (*SeqColDict*): Sequence collection A
- `B` (*SeqColDict*): Sequence collection B

**Returns:**

- *dict*: Following formal seqcol specification comparison function return value

<a id="refget.utils.calc_jaccard_similarities"></a>

### `calc_jaccard_similarities`

```python
def calc_jaccard_similarities(A: SeqColDict, B: SeqColDict) -> dict[str, float]
```

Takes two sequence collections and calculates jaccard similarties for all attributes

**Parameters:**

- `A` (*SeqColDict*): Sequence collection A
- `B` (*SeqColDict*): Sequence collection B

**Returns:**

- `dict` (*dict[str, float]*): Jaccard similarities for all attributes

<a id="refget.utils.validate_seqcol"></a>

### `validate_seqcol`

```python
def validate_seqcol(seqcol_obj: SeqColDict, schema=None) -> bool
```

Validate a seqcol object against the seqcol schema.
Returns True if valid, raises InvalidSeqColError if not, which enumerates the errors.
Retrieve individual errors with exception.errors

<a id="refget.utils.validate_seqcol_bool"></a>

### `validate_seqcol_bool`

```python
def validate_seqcol_bool(seqcol_obj: SeqColDict, schema=None) -> bool
```

Validate a seqcol object against the seqcol schema. Returns True if valid, False if not.

To enumerate the errors, use validate_seqcol instead.

### FastAPI Integration

<a id="refget.router.create_refget_router"></a>

### `create_refget_router`

```python
def create_refget_router(sequences: bool = False, collections: bool = True, pangenomes: bool = False, fasta_drs: bool = False, compliance: bool = True, refget_store_url: str = None, mount_prefix: str = '') -> APIRouter
```

Create a FastAPI router for the sequence collection API.
This router provides endpoints for retrieving and comparing sequence collections.
You can choose which endpoints to include by setting the sequences, collections,
pangenomes, or fasta_drs flags.

**Parameters:**

- `sequences` (*bool*): Include sequence endpoints (default: `False`)
- `collections` (*bool*): Include sequence collection endpoints (default: `True`)
- `pangenomes` (*bool*): Include pangenome endpoints (default: `False`)
- `fasta_drs` (*bool*): Include FASTA DRS endpoints (default: `False`)
- `refget_store_url` (*str*): URL of backing RefgetStore (e.g., s3://bucket/store/) (default: `None`)
- `mount_prefix` (*str*): The path prefix this router will be included under,
  when it is included with `include_router(..., prefix=...)` rather
  than mounted as a sub-application. Only used so the compliance
  endpoints self-target the seqcol service instead of the server root.
  Leave empty when mounting an app (`scope["root_path"]` covers it). (default: `''`)

**Returns:**

- *APIRouter*: A FastAPI router with the specified endpoints

**Examples:**

```
app.include_router(create_refget_router(fasta_drs=True), prefix="/seqcol")
```

## Client Classes

The client module provides interfaces for interacting with refget-compliant servers.

<a id="refget.clients.SequenceClient"></a>

### `SequenceClient`

```python
class SequenceClient(urls: list[str] = ['https://www.ebi.ac.uk/ena/cram'], raise_errors: Optional[bool] = None)
```

Bases: `RefgetClient`

A client for interacting with a refget sequences API.

Initializes the sequences client.

**Parameters:**

- `urls` (*list*): A list of base URLs of the sequences API. Defaults to ["https://www.ebi.ac.uk/ena/cram/sequence/"]. (default: `['https://www.ebi.ac.uk/ena/cram']`)
- `raise_errors` (*bool*): Whether to raise errors or log them. Defaults to None, which will guess. (default: `None`)

Attributes:
    urls (list): The list of base URLs of the sequences API.

#### Attributes

- <a id="refget.clients.SequenceClient.raise_errors"></a>**`raise_errors`**
- <a id="refget.clients.SequenceClient.urls"></a>**`urls`**

#### Methods

<a id="refget.clients.SequenceClient.get_metadata"></a>

##### `get_metadata`

```python
def get_metadata(digest: str) -> Optional[dict]
```

Retrieves metadata for a given sequence digest.

**Parameters:**

- `digest` (*str*): The digest of the sequence.

**Returns:**

- *dict*: The metadata.

<a id="refget.clients.SequenceClient.get_sequence"></a>

##### `get_sequence`

```python
def get_sequence(digest: str, start: Optional[int] = None, end: Optional[int] = None) -> Optional[str]
```

Retrieves a sequence for a given digest.

**Parameters:**

- `digest` (*str*): The digest of the sequence.

**Returns:**

- *str*: The sequence.

<a id="refget.clients.SequenceCollectionClient"></a>

### `SequenceCollectionClient`

```python
class SequenceCollectionClient(urls: list[str] = ['https://seqcolapi.databio.org'], raise_errors: Optional[bool] = None)
```

Bases: `RefgetClient`

A client for interacting with a refget sequence collections API.

Initializes the sequence collection client.

**Parameters:**

- `urls` (*list*): A list of base URLs of the sequence collection API. Defaults to ["https://seqcolapi.databio.org"]. (default: `['https://seqcolapi.databio.org']`)

**Attributes:**

- `urls` (*list*): The list of base URLs of the sequence collection API.

#### Attributes

- <a id="refget.clients.SequenceCollectionClient.raise_errors"></a>**`raise_errors`**
- <a id="refget.clients.SequenceCollectionClient.urls"></a>**`urls`**

#### Methods

<a id="refget.clients.SequenceCollectionClient.aliases_for"></a>

##### `aliases_for`

```python
def aliases_for(digest: str, kind: str = 'collection') -> Optional[dict]
```

Reverse lookup: list all (namespace, alias) pairs for a digest.

**Parameters:**

- `digest` (*str*): The digest to look up.
- `kind` (*str*): "collection" (default) or "sequence". (default: `'collection'`)

**Returns:**

- *dict*: {"digest": ..., "aliases": [[namespace, alias], ...]}.

<a id="refget.clients.SequenceCollectionClient.build_chrom_sizes"></a>

##### `build_chrom_sizes`

```python
def build_chrom_sizes(digest: str) -> str
```

Build a chrom.sizes file content for a sequence collection.

Format per line: NAME\tLENGTH

**Parameters:**

- `digest` (*str*): The sequence collection digest

**Returns:**

- *str*: String content of the chrom.sizes file

<a id="refget.clients.SequenceCollectionClient.build_fai"></a>

##### `build_fai`

```python
def build_fai(digest: str) -> str
```

Build a complete .fai index file content for a FASTA.

FAI format per line: NAME\tLENGTH\tOFFSET\tLINEBASES\tLINEWIDTH

**Parameters:**

- `digest` (*str*): The sequence collection digest

**Returns:**

- *str*: String content of the .fai file

<a id="refget.clients.SequenceCollectionClient.compare"></a>

##### `compare`

```python
def compare(digest1: str, digest2: str) -> Optional[dict]
```

Compares two sequence collections hosted on the server.

**Parameters:**

- `digest1` (*str*): The digest of the first sequence collection.
- `digest2` (*str*): The digest of the second sequence collection.

**Returns:**

- *dict*: The JSON response containing the comparison of the two sequence collections.

<a id="refget.clients.SequenceCollectionClient.compare_local"></a>

##### `compare_local`

```python
def compare_local(digest: str, local_collection: dict) -> Optional[dict]
```

Compares a server-hosted sequence collection with a local collection.

**Parameters:**

- `digest` (*str*): The digest of the server-hosted sequence collection.
- `local_collection` (*dict*): A level 2 sequence collection representation.

**Returns:**

- *dict*: The JSON response containing the comparison.

<a id="refget.clients.SequenceCollectionClient.download_fasta"></a>

##### `download_fasta`

```python
def download_fasta(digest: str, dest_path: str = None, access_id: str = None) -> str
```

Download the FASTA file to a local path.

**Parameters:**

- `digest` (*str*): The sequence collection digest
- `dest_path` (*str*): Destination file path. If None, uses object name. (default: `None`)
- `access_id` (*str*): Specific access method to use. If None, tries all. (default: `None`)

**Returns:**

- *str*: Path to downloaded file

**Raises:**

- *ValueError*: If no access methods available or specified access_id not found

<a id="refget.clients.SequenceCollectionClient.download_fasta_to_store"></a>

##### `download_fasta_to_store`

```python
def download_fasta_to_store(digest: str, store: 'RefgetStore', access_id: str = None, temp_dir: str = None, namespaces: Optional[list[str]] = None) -> str
```

Download the FASTA file and import it into a RefgetStore.

This method downloads the FASTA file from the DRS endpoint and immediately
imports it into the provided RefgetStore, enabling local sequence retrieval
by digest without re-downloading.

**Parameters:**

- `digest` (*str*): The sequence collection digest
- `store` (*RefgetStore*): The RefgetStore instance to import into
- `access_id` (*str*): Specific access method to use. If None, tries all. (default: `None`)
- `temp_dir` (*str*): Directory for temporary download. If None, uses system temp. (default: `None`)

**Returns:**

- *str*: The collection digest of the imported sequences

**Raises:**

- *ValueError*: If no access methods available or specified access_id not found
- *ImportError*: If gtars/RefgetStore is not available

:::note[Example]
```python
>>> from refget.store import RefgetStore
>>> from refget.clients import SequenceCollectionClient
>>> store = RefgetStore.in_memory()
>>> client = SequenceCollectionClient()
>>> collection_digest = client.download_fasta_to_store("abc123", store)
>>> # Now you can retrieve sequences by digest from the local store
>>> seq = store.get_substring(sequence_digest, 0, 100)
```
:::

<a id="refget.clients.SequenceCollectionClient.get_attribute"></a>

##### `get_attribute`

```python
def get_attribute(attribute: str, digest: str) -> Optional[dict]
```

Retrieves a specific attribute value by its digest.

**Parameters:**

- `attribute` (*str*): The attribute name (e.g., "names", "lengths", "sequences").
- `digest` (*str*): The level 1 digest of the attribute.

**Returns:**

- *dict*: The JSON response containing the attribute value.

<a id="refget.clients.SequenceCollectionClient.get_collection"></a>

##### `get_collection`

```python
def get_collection(digest: str, level: int = 2) -> Optional[dict]
```

Retrieves a sequence collection for a given digest and detail level.

**Parameters:**

- `digest` (*str*): The digest of the sequence collection.
- `level` (*int*): The level of detail for the sequence collection. Defaults to 2. (default: `2`)

**Returns:**

- *dict*: The JSON response containing the sequence collection.

<a id="refget.clients.SequenceCollectionClient.get_fasta"></a>

##### `get_fasta`

```python
def get_fasta(digest: str) -> Optional[dict]
```

Get DRS object metadata for a FASTA file.

**Parameters:**

- `digest` (*str*): The sequence collection digest (which is also the DRS object ID)

**Returns:**

- *dict*: DRS object with id, self_uri, size, checksums, access_methods, etc.

<a id="refget.clients.SequenceCollectionClient.get_fasta_index"></a>

##### `get_fasta_index`

```python
def get_fasta_index(digest: str) -> Optional[dict]
```

Get FAI index data for a FASTA file.

**Parameters:**

- `digest` (*str*): The sequence collection digest

**Returns:**

- *dict*: Dict with line_bases, extra_line_bytes, offsets

<a id="refget.clients.SequenceCollectionClient.get_fhr"></a>

##### `get_fhr`

```python
def get_fhr(digest: str) -> Optional[dict]
```

Get FHR metadata for a collection.

**Parameters:**

- `digest` (*str*): The collection digest.

**Returns:**

- *dict*: FHR metadata, or None if not found.

<a id="refget.clients.SequenceCollectionClient.get_refget_store"></a>

##### `get_refget_store`

```python
def get_refget_store(cache_dir: str) -> 'RefgetStore'
```

Get a RefgetStore instance connected to the server's backing store.

**Parameters:**

- `cache_dir` (*str*): Local directory for caching store data

**Returns:**

- *RefgetStore*: RefgetStore instance loaded from remote

**Raises:**

- *ValueError*: If server doesn't have a RefgetStore configured
- *ImportError*: If gtars is not installed

<a id="refget.clients.SequenceCollectionClient.get_refget_store_url"></a>

##### `get_refget_store_url`

```python
def get_refget_store_url() -> Optional[str]
```

Discover RefgetStore URL from service-info if available.

**Returns:**

- *str*: The RefgetStore URL if configured, None otherwise.

<a id="refget.clients.SequenceCollectionClient.get_regions"></a>

##### `get_regions`

```python
def get_regions(digest: str, regions: list) -> Optional[list]
```

Extract region substrings from a server-hosted sequence collection.

Posts a list of regions to the server's region-extraction endpoint and
returns the structured results. Requires a store-backed server (the
database backend responds with HTTP 501).

**Parameters:**

- `digest` (*str*): The collection digest to extract regions from.
- `regions` (*list*): A list of {"chrom", "start", "end"} dicts.

**Returns:**

- *list*: A list of {"chrom_name", "start", "end", "sequence"} dicts.

<a id="refget.clients.SequenceCollectionClient.is_aliases_enabled"></a>

##### `is_aliases_enabled`

```python
def is_aliases_enabled() -> bool
```

Check if alias endpoints are advertised in service-info.

**Returns:**

- *bool*: True if aliases are enabled, False otherwise.

<a id="refget.clients.SequenceCollectionClient.is_fasta_drs_enabled"></a>

##### `is_fasta_drs_enabled`

```python
def is_fasta_drs_enabled() -> bool
```

Check if FastaDRS endpoints are available.

**Returns:**

- *bool*: True if FastaDRS is enabled, False otherwise.

<a id="refget.clients.SequenceCollectionClient.is_fhr_enabled"></a>

##### `is_fhr_enabled`

```python
def is_fhr_enabled() -> bool
```

Check if FHR metadata is advertised in service-info.

**Returns:**

- *bool*: True if FHR metadata is enabled, False otherwise.

<a id="refget.clients.SequenceCollectionClient.list_alias_namespaces"></a>

##### `list_alias_namespaces`

```python
def list_alias_namespaces(kind: str = 'collection') -> Optional[dict]
```

List alias namespaces for the given kind.

**Parameters:**

- `kind` (*str*): "collection" (default) or "sequence". (default: `'collection'`)

**Returns:**

- *dict*: {"namespaces": [...]}.

<a id="refget.clients.SequenceCollectionClient.list_aliases"></a>

##### `list_aliases`

```python
def list_aliases(namespace: str, kind: str = 'collection') -> Optional[dict]
```

List aliases within a namespace.

**Parameters:**

- `namespace` (*str*): The alias namespace.
- `kind` (*str*): "collection" (default) or "sequence". (default: `'collection'`)

**Returns:**

- *dict*: {"namespace": ..., "aliases": [...]}.

<a id="refget.clients.SequenceCollectionClient.list_attributes"></a>

##### `list_attributes`

```python
def list_attributes(attribute: str, page: Optional[int] = None, page_size: Optional[int] = None) -> Optional[dict]
```

Lists all available values for a given attribute with optional paging support.

**Parameters:**

- `attribute` (*str*): The attribute to list values for.
- `page` (*int*): The page number to retrieve. Defaults to None. (default: `None`)
- `page_size` (*int*): The number of items per page. Defaults to None. (default: `None`)

**Returns:**

- *dict*: The JSON response containing the list of available values for the attribute.

<a id="refget.clients.SequenceCollectionClient.list_collections"></a>

##### `list_collections`

```python
def list_collections(page: Optional[int] = None, page_size: Optional[int] = None, **filters) -> Optional[dict]
```

Lists all available sequence collections with optional paging and attribute filtering support.

**Parameters:**

- `page` (*int*): The page number to retrieve. Defaults to None. (default: `None`)
- `page_size` (*int*): The number of items per page. Defaults to None. (default: `None`)
- `**filters` (*Any*): Optional attribute filters (e.g., names="abc123", lengths="def456").
        Values should be level 1 digests of the attributes. (default: `{}`)

**Returns:**

- *dict*: The JSON response containing the list of available sequence collections.

<a id="refget.clients.SequenceCollectionClient.list_fhr"></a>

##### `list_fhr`

```python
def list_fhr() -> Optional[dict]
```

List collections that have FHR metadata.

**Returns:**

- *dict*: {"collections": [...]}.

<a id="refget.clients.SequenceCollectionClient.resolve_alias"></a>

##### `resolve_alias`

```python
def resolve_alias(namespace: str, alias: str, kind: str = 'collection') -> Optional[dict]
```

Resolve a namespace:alias to a digest.

**Parameters:**

- `namespace` (*str*): The alias namespace.
- `alias` (*str*): The alias name.
- `kind` (*str*): "collection" (default) or "sequence". (default: `'collection'`)

**Returns:**

- *dict*: {"namespace": ..., "alias": ..., "digest": ...} or None.

<a id="refget.clients.SequenceCollectionClient.service_info"></a>

##### `service_info`

```python
def service_info() -> Optional[dict]
```

Retrieves information about the service.

**Returns:**

- *dict*: The service information.

<a id="refget.clients.SequenceCollectionClient.write_chrom_sizes"></a>

##### `write_chrom_sizes`

```python
def write_chrom_sizes(digest: str, dest_path: str) -> str
```

Write a chrom.sizes file for a sequence collection.

**Parameters:**

- `digest` (*str*): The sequence collection digest
- `dest_path` (*str*): Path to write the chrom.sizes file

**Returns:**

- *str*: Path to the written file

<a id="refget.clients.SequenceCollectionClient.write_fai"></a>

##### `write_fai`

```python
def write_fai(digest: str, dest_path: str) -> str
```

Write a .fai index file for a FASTA.

**Parameters:**

- `digest` (*str*): The sequence collection digest
- `dest_path` (*str*): Path to write the .fai file

**Returns:**

- *str*: Path to the written file

<a id="refget.clients.FastaDrsClient"></a>

### `FastaDrsClient`

```python
class FastaDrsClient(urls: list[str] = ['https://seqcolapi.databio.org/fasta'], raise_errors: Optional[bool] = None)
```

Bases: `RefgetClient`

A client for interacting with FASTA files via GA4GH DRS endpoints.

Initializes the FASTA DRS client.

**Parameters:**

- `urls` (*list*): A list of base URLs of the FASTA DRS API.
  Defaults to ["https://seqcolapi.databio.org/fasta"]. (default: `['https://seqcolapi.databio.org/fasta']`)
- `raise_errors` (*bool*): Whether to raise errors or log them.
  Defaults to None, which will guess. (default: `None`)

**Attributes:**

- `urls` (*list*): The list of base URLs of the FASTA DRS API.

#### Attributes

- <a id="refget.clients.FastaDrsClient.raise_errors"></a>**`raise_errors`**
- <a id="refget.clients.FastaDrsClient.urls"></a>**`urls`**

#### Methods

<a id="refget.clients.FastaDrsClient.build_fai"></a>

##### `build_fai`

```python
def build_fai(digest: str, seqcol_client: 'SequenceCollectionClient' = None) -> str
```

Build a complete .fai index file content for a FASTA.

FAI format per line: NAME       LENGTH  OFFSET  LINEBASES       LINEWIDTH

**Parameters:**

- `digest` (*str*): The sequence collection digest
- `seqcol_client` (*[SequenceCollectionClient](#refget.clients.SequenceCollectionClient)*): SequenceCollectionClient
  to use. If None, uses parent client or creates one. (default: `None`)

**Returns:**

- *str*: String content of the .fai file

<a id="refget.clients.FastaDrsClient.download"></a>

##### `download`

```python
def download(digest: str, dest_path: str = None, access_id: str = None) -> str
```

Download the FASTA file to a local path.

**Parameters:**

- `digest` (*str*): The sequence collection digest
- `dest_path` (*str*): Destination file path. If None, uses object name. (default: `None`)
- `access_id` (*str*): Specific access method to use. If None, tries all. (default: `None`)

**Returns:**

- *str*: Path to downloaded file

**Raises:**

- *ValueError*: If no access methods available or specified access_id not found

<a id="refget.clients.FastaDrsClient.download_to_store"></a>

##### `download_to_store`

```python
def download_to_store(digest: str, store: 'RefgetStore', access_id: str = None, temp_dir: str = None, namespaces: Optional[list[str]] = None) -> str
```

Download the FASTA file and import it into a RefgetStore.

This method downloads the FASTA file from the DRS endpoint and immediately
imports it into the provided RefgetStore, enabling local sequence retrieval
by digest without re-downloading.

**Parameters:**

- `digest` (*str*): The sequence collection digest
- `store` (*RefgetStore*): The RefgetStore instance to import into
- `access_id` (*str*): Specific access method to use. If None, tries all. (default: `None`)
- `temp_dir` (*str*): Directory for temporary download. If None, uses system temp. (default: `None`)
- `namespaces` (*list[str]*): Namespace prefixes to extract aliases from
  FASTA headers when importing into the store. (default: `None`)

**Returns:**

- *str*: The collection digest of the imported sequences

**Raises:**

- *ValueError*: If no access methods available or specified access_id not found
- *ImportError*: If gtars/RefgetStore is not available

:::note[Example]
```python
>>> from refget.store import RefgetStore
>>> store = RefgetStore.in_memory()
>>> client = FastaDrsClient()
>>> collection_digest = client.download_to_store("abc123", store)
```
:::

<a id="refget.clients.FastaDrsClient.get_access_url"></a>

##### `get_access_url`

```python
def get_access_url(digest: str, access_id: str) -> Optional[dict]
```

Get access URL for a specific access method.

**Parameters:**

- `digest` (*str*): The sequence collection digest
- `access_id` (*str*): The access ID from the access method

**Returns:**

- *dict*: Access URL object

<a id="refget.clients.FastaDrsClient.get_index"></a>

##### `get_index`

```python
def get_index(digest: str) -> Optional[dict]
```

Get FAI index data for a FASTA file.

**Parameters:**

- `digest` (*str*): The sequence collection digest

**Returns:**

- *dict*: Dict with line_bases, extra_line_bytes, offsets

<a id="refget.clients.FastaDrsClient.get_object"></a>

##### `get_object`

```python
def get_object(digest: str) -> Optional[dict]
```

Get DRS object metadata for a FASTA file.

**Parameters:**

- `digest` (*str*): The sequence collection digest (which is also the DRS object ID)

**Returns:**

- *dict*: DRS object with id, self_uri, size, checksums, access_methods, etc.

<a id="refget.clients.FastaDrsClient.service_info"></a>

##### `service_info`

```python
def service_info() -> Optional[dict]
```

Get DRS service info.

**Returns:**

- *dict*: The service information.

<a id="refget.clients.FastaDrsClient.write_fai"></a>

##### `write_fai`

```python
def write_fai(digest: str, dest_path: str, seqcol_client: 'SequenceCollectionClient' = None) -> str
```

Write a .fai index file for a FASTA.

**Parameters:**

- `digest` (*str*): The sequence collection digest
- `dest_path` (*str*): Path to write the .fai file
- `seqcol_client` (*[SequenceCollectionClient](#refget.clients.SequenceCollectionClient)*): SequenceCollectionClient to use (default: `None`)

**Returns:**

- *str*: Path to the written file

<a id="refget.clients.PangenomeClient"></a>

### `PangenomeClient`

```python
class PangenomeClient
```

Bases: `RefgetClient`

## Agent Classes

Agents provide higher-level abstractions for working with refget data in a PostgreSQL database.

<a id="refget.agents.RefgetDBAgent"></a>

### `RefgetDBAgent`

```python
class RefgetDBAgent(engine: Optional[SqlalchemyDatabaseEngine] = None, postgres_str: Optional[str] = None, schema=SEQCOL_SCHEMA_PATH, inherent_attrs: List[str] = DEFAULT_INHERENT_ATTRS, fasta_drs_url_prefix: Optional[str] = None)
```

Primary aggregator agent, interface to all other agents

Parameterized it via these environment variables:
- POSTGRES_HOST
- POSTGRES_DB
- POSTGRES_USER
- POSTGRES_PASSWORD

#### Attributes

- <a id="refget.agents.RefgetDBAgent.engine"></a>**`engine`**
- <a id="refget.agents.RefgetDBAgent.inherent_attrs"></a>**`inherent_attrs`**
- <a id="refget.agents.RefgetDBAgent.schema_dict"></a>**`schema_dict`**

#### Properties

<a id="refget.agents.RefgetDBAgent.attribute"></a>

##### `attribute`

```python
@property
def attribute -> AttributeAgent
```

<a id="refget.agents.RefgetDBAgent.fasta_drs"></a>

##### `fasta_drs`

```python
@property
def fasta_drs -> FastaDrsAgent
```

<a id="refget.agents.RefgetDBAgent.pangenome"></a>

##### `pangenome`

```python
@property
def pangenome -> PangenomeAgent
```

<a id="refget.agents.RefgetDBAgent.seq"></a>

##### `seq`

```python
@property
def seq -> SequenceAgent
```

<a id="refget.agents.RefgetDBAgent.seqcol"></a>

##### `seqcol`

```python
@property
def seqcol -> SequenceCollectionAgent
```

#### Methods

<a id="refget.agents.RefgetDBAgent.aliases_for"></a>

##### `aliases_for`

```python
def aliases_for(kind: str, digest: str) -> list
```

<a id="refget.agents.RefgetDBAgent.calc_similarities"></a>

##### `calc_similarities`

```python
def calc_similarities(digestA: str, digestB: str) -> dict
```

Calculates the Jaccard similarity between two sequence collections.

This method retrieves two sequence collections using their digests and then
computes jaccard similarities for all attributes.

**Parameters:**

- `digestA` (*str*): The digest (identifier) for the first sequence collection.
- `digestB` (*str*): The digest (identifier) for the second sequence collection.

**Returns:**

- *dict*: The Jaccard similarity score between the two sequence collections for all present and shared attributes.

<a id="refget.agents.RefgetDBAgent.calc_similarities_seqcol_dicts"></a>

##### `calc_similarities_seqcol_dicts`

```python
def calc_similarities_seqcol_dicts(seqcolA: dict, seqcolB: dict) -> dict
```

Calculates the Jaccard similarity between two sequence collections.

This method retrieves one sequence collections using a digests and then
computes jaccard similarities versus another input sequence collection dictionary.

**Parameters:**

- `seqcolA` (*dict*): the first sequence collection in dict format.
- `seqcolB` (*dict*): the second sequence collection in dict format.

**Returns:**

- *dict*: The Jaccard similarity score between the two sequence collections for all present and shared attributes.

<a id="refget.agents.RefgetDBAgent.capabilities"></a>

##### `capabilities`

```python
def capabilities() -> dict
```

<a id="refget.agents.RefgetDBAgent.collection_count"></a>

##### `collection_count`

```python
def collection_count() -> int
```

<a id="refget.agents.RefgetDBAgent.compare_1_digest"></a>

##### `compare_1_digest`

```python
def compare_1_digest(digestA: str, seqcolB: dict) -> dict
```

<a id="refget.agents.RefgetDBAgent.compare_digest_with_level2"></a>

##### `compare_digest_with_level2`

```python
def compare_digest_with_level2(digest: str, level2_b: dict) -> dict
```

<a id="refget.agents.RefgetDBAgent.compare_digests"></a>

##### `compare_digests`

```python
def compare_digests(digestA: str, digestB: str) -> dict
```

<a id="refget.agents.RefgetDBAgent.get_attribute"></a>

##### `get_attribute`

```python
def get_attribute(attribute_name: str, attribute_digest: str) -> list
```

<a id="refget.agents.RefgetDBAgent.get_collection"></a>

##### `get_collection`

```python
def get_collection(digest: str, level: int = 2) -> dict
```

<a id="refget.agents.RefgetDBAgent.get_collection_attribute"></a>

##### `get_collection_attribute`

```python
def get_collection_attribute(digest: str, attribute: str) -> list
```

<a id="refget.agents.RefgetDBAgent.get_collection_itemwise"></a>

##### `get_collection_itemwise`

```python
def get_collection_itemwise(digest: str, limit: int | None = None) -> list[dict]
```

<a id="refget.agents.RefgetDBAgent.get_fhr"></a>

##### `get_fhr`

```python
def get_fhr(digest: str)
```

<a id="refget.agents.RefgetDBAgent.list_alias_namespaces"></a>

##### `list_alias_namespaces`

```python
def list_alias_namespaces(kind: str) -> list
```

<a id="refget.agents.RefgetDBAgent.list_aliases"></a>

##### `list_aliases`

```python
def list_aliases(kind: str, namespace: str) -> list
```

<a id="refget.agents.RefgetDBAgent.list_attributes"></a>

##### `list_attributes`

```python
def list_attributes(attribute: str, page: int = 0, page_size: int = 100) -> dict
```

<a id="refget.agents.RefgetDBAgent.list_collections"></a>

##### `list_collections`

```python
def list_collections(page: int = 0, page_size: int = 100, filters: dict | None = None) -> dict
```

<a id="refget.agents.RefgetDBAgent.list_fhr"></a>

##### `list_fhr`

```python
def list_fhr() -> list
```

<a id="refget.agents.RefgetDBAgent.resolve_alias"></a>

##### `resolve_alias`

```python
def resolve_alias(kind: str, namespace: str, alias: str)
```

<a id="refget.agents.RefgetDBAgent.retrieve_level2_digest"></a>

##### `retrieve_level2_digest`

```python
def retrieve_level2_digest(seqcoldigest: str) -> dict
```

<a id="refget.agents.RefgetDBAgent.truncate"></a>

##### `truncate`

```python
def truncate() -> int
```

Delete all records from the database

<a id="refget.agents.SequenceCollectionAgent"></a>

### `SequenceCollectionAgent`

```python
class SequenceCollectionAgent(engine: SqlalchemyDatabaseEngine, inherent_attrs: Optional[List[str]] = None, parent: Optional['RefgetDBAgent'] = None)
```

Agent for interacting with database of sequence collection

#### Attributes

- <a id="refget.agents.SequenceCollectionAgent.engine"></a>**`engine`**
- <a id="refget.agents.SequenceCollectionAgent.inherent_attrs"></a>**`inherent_attrs`**
- <a id="refget.agents.SequenceCollectionAgent.parent"></a>**`parent`**

#### Methods

<a id="refget.agents.SequenceCollectionAgent.add"></a>

##### `add`

```python
def add(seqcol: SequenceCollection, update: bool = False) -> SequenceCollection
```

Add a sequence collection to the database or update it if it exists

**Parameters:**

- `seqcol` (*SequenceCollection*): The sequence collection to add
- `update` (*bool*): If True, update an existing collection if it exists (default: `False`)

**Returns:**

- *SequenceCollection*: The added or updated sequence collection

<a id="refget.agents.SequenceCollectionAgent.add_from_dict"></a>

##### `add_from_dict`

```python
def add_from_dict(seqcol_dict: dict, update: bool = False) -> SequenceCollection
```

Add a sequence collection from a seqcol dictionary

**Parameters:**

- `seqcol_dict` (*dict*): The sequence collection in dictionary form
- `update` (*bool*): If True, update an existing collection if it exists (default: `False`)

**Returns:**

- *SequenceCollection*: The added or updated sequence collection

<a id="refget.agents.SequenceCollectionAgent.add_from_fasta_file"></a>

##### `add_from_fasta_file`

```python
def add_from_fasta_file(fasta_file_path: str, update: bool = False, create_fasta_drs: bool = True, human_readable_name: str = None) -> SequenceCollection
```

Given a path to a fasta file, load the sequences into the refget database.

**Parameters:**

- `fasta_file_path` (*str*): Path to the fasta file
- `update` (*bool*): If True, update an existing collection if it exists (default: `False`)
- `create_fasta_drs` (*bool*): If True, create a FastaDrsObject for the FASTA file (default: `True`)
- `human_readable_name` (*str*): Optional human-readable name for the collection (default: `None`)

**Returns:**

- *SequenceCollection*: The added or updated sequence collection

<a id="refget.agents.SequenceCollectionAgent.add_from_fasta_file_with_name"></a>

##### `add_from_fasta_file_with_name`

```python
def add_from_fasta_file_with_name(fasta_file_path: str, human_readable_name: str, update: bool = False, create_fasta_drs: bool = True) -> SequenceCollection
```

Given a path to a fasta file, and a human-readable name, load the sequences into the refget database.

Deprecated: Use add_from_fasta_file(fasta_file_path, human_readable_name=name) instead.

<a id="refget.agents.SequenceCollectionAgent.add_from_fasta_pep"></a>

##### `add_from_fasta_pep`

```python
def add_from_fasta_pep(pep: 'peppy.Project', fa_root: str, update: bool = False, create_fasta_drs: bool = True) -> dict
```

Given a PEP project and a root directory containing the fasta files,
load the fasta files into the refget database.

**Parameters:**

- `pep` (*peppy.Project*): PEP project object containing sample metadata
- `fa_root` (*str*): Root directory containing the fasta files
- `update` (*bool*): If True, update existing sequence collections (default: `False`)
- `create_fasta_drs` (*bool*): If True, create FastaDrsObjects for the FASTA files (default: `True`)

**Returns:**

- *dict*: A dictionary of the digests of the added sequence collections

<a id="refget.agents.SequenceCollectionAgent.get"></a>

##### `get`

```python
def get(digest: str, return_format: str = 'level2', attribute: Optional[str] = None, itemwise_limit: Optional[int] = None) -> SequenceCollection | dict | list
```

Get a sequence collection by digest

**Parameters:**

- `digest` (*str*): The digest of the sequence collection
- `return_format` (*str*): The format in which to return the sequence collection (default: `'level2'`)
- `attribute` (*str*): Name of an attribute to return, if you just want an attribute (default: `None`)
- `itemwise_limit` (*int*): Limit the number of items returned in itemwise format (default: `None`)

**Returns:**

- *SequenceCollection*: The sequence collection (in requested format)

<a id="refget.agents.SequenceCollectionAgent.get_many_level2_offset"></a>

##### `get_many_level2_offset`

```python
def get_many_level2_offset(limit: int = 50, offset: int = 0, target_digests: Optional[List[str]] = None) -> ResultsSequenceCollections
```

<a id="refget.agents.SequenceCollectionAgent.list"></a>

##### `list`

```python
def list(page_size: int = 100, cursor: Optional[str] = None) -> dict
```

<a id="refget.agents.SequenceCollectionAgent.list_by_offset"></a>

##### `list_by_offset`

```python
def list_by_offset(limit: int = 50, offset: int = 0) -> dict
```

<a id="refget.agents.SequenceCollectionAgent.search_by_attributes"></a>

##### `search_by_attributes`

```python
def search_by_attributes(filters: dict, offset: int = 0, limit: int = 50) -> dict
```

Search sequence collections by multiple attribute filters (AND logic).

**Parameters:**

- `filters` (*dict*): Dict of {attribute_name: digest} pairs
- `offset` (*int*): Pagination offset (default: `0`)
- `limit` (*int*): Max results to return (default: `50`)

**Returns:**

- *dict*: Dict with pagination info and results

<a id="refget.agents.SequenceAgent"></a>

### `SequenceAgent`

```python
class SequenceAgent(engine: SqlalchemyDatabaseEngine)
```

Agent for interacting with database of sequences

#### Attributes

- <a id="refget.agents.SequenceAgent.engine"></a>**`engine`**

#### Methods

<a id="refget.agents.SequenceAgent.add"></a>

##### `add`

```python
def add(sequence: Sequence) -> Sequence
```

<a id="refget.agents.SequenceAgent.get"></a>

##### `get`

```python
def get(digest: str, start: int | None = None, end: int | None = None) -> str
```

<a id="refget.agents.SequenceAgent.list"></a>

##### `list`

```python
def list(offset: int = 0, limit: int = 50) -> dict
```

<a id="refget.agents.PangenomeAgent"></a>

### `PangenomeAgent`

```python
class PangenomeAgent(parent: 'RefgetDBAgent')
```

Agent for interacting with database of pangenomes

#### Attributes

- <a id="refget.agents.PangenomeAgent.engine"></a>**`engine`**
- <a id="refget.agents.PangenomeAgent.parent"></a>**`parent`**

#### Methods

<a id="refget.agents.PangenomeAgent.add"></a>

##### `add`

```python
def add(pangenome: Pangenome) -> Pangenome
```

<a id="refget.agents.PangenomeAgent.add_from_fasta_pep"></a>

##### `add_from_fasta_pep`

```python
def add_from_fasta_pep(pep: 'peppy.Project', fa_root: str, update: bool = False) -> Pangenome
```

<a id="refget.agents.PangenomeAgent.get"></a>

##### `get`

```python
def get(digest: str, return_format: str = 'level2') -> Pangenome | dict
```

<a id="refget.agents.PangenomeAgent.list_by_offset"></a>

##### `list_by_offset`

```python
def list_by_offset(limit: int = 50, offset: int = 0) -> dict
```

<a id="refget.agents.AttributeAgent"></a>

### `AttributeAgent`

```python
class AttributeAgent(engine: SqlalchemyDatabaseEngine)
```

#### Attributes

- <a id="refget.agents.AttributeAgent.engine"></a>**`engine`**

#### Methods

<a id="refget.agents.AttributeAgent.get"></a>

##### `get`

```python
def get(attribute_type: str, digest: str) -> list
```

<a id="refget.agents.AttributeAgent.list"></a>

##### `list`

```python
def list(attribute_type: str, offset: int = 0, limit: int = 50) -> dict
```

<a id="refget.agents.AttributeAgent.search"></a>

##### `search`

```python
def search(attribute_type: str, digest: str, offset: int = 0, limit: int = 50) -> dict
```

<a id="refget.agents.FastaDrsAgent"></a>

### `FastaDrsAgent`

```python
class FastaDrsAgent(engine: SqlalchemyDatabaseEngine, url_prefix: Optional[str] = None)
```

Agent for interacting with database of FASTA DRS objects

#### Attributes

- <a id="refget.agents.FastaDrsAgent.engine"></a>**`engine`**
- <a id="refget.agents.FastaDrsAgent.url_prefix"></a>**`url_prefix`**

#### Methods

<a id="refget.agents.FastaDrsAgent.add"></a>

##### `add`

```python
def add(fasta_drs: FastaDrsObject) -> FastaDrsObject
```

Add a FastaDrsObject to the database

<a id="refget.agents.FastaDrsAgent.add_access_method"></a>

##### `add_access_method`

```python
def add_access_method(digest: str, access_method: AccessMethod) -> FastaDrsObject
```

Add an access method to an existing FastaDrsObject.

**Parameters:**

- `digest` (*str*): The digest (object_id) of the DRS object
- `access_method` (*AccessMethod*): The AccessMethod to add

**Returns:**

- *FastaDrsObject*: The updated FastaDrsObject

<a id="refget.agents.FastaDrsAgent.get"></a>

##### `get`

```python
def get(digest: str) -> FastaDrsObject
```

Get a FastaDrsObject by its digest (object_id)

<a id="refget.agents.FastaDrsAgent.list_by_offset"></a>

##### `list_by_offset`

```python
def list_by_offset(limit: int = 50, offset: int = 0) -> dict
```

List FastaDrsObjects with pagination

## RefgetStore (gtars)

RefgetStore provides high-performance local sequence storage implemented in Rust. It supports:

- **In-memory and on-disk storage** with optional compression
- **Remote store access** with local caching
- **Sequence retrieval** by digest or by collection + name
- **BED file region extraction** for batch operations
- **FASTA export** for individual sequences or regions

See the [RefgetStore tutorial](/refget/using-services/refgetstore.md) for usage examples.

`RefgetStore`, `digest_fasta`, `compute_fai`, `digest_sequence`, `SequenceCollection`, `StorageMode`, and `sha512t24u_digest` are implemented in Rust in [gtars](https://docs.bedbase.org/gtars/refget/) and re-exported by refget. Their API reference is in the gtars docs: [gtars refget API reference](https://docs.bedbase.org/gtars/python/refget-api/).

## Digest Functions

Low-level functions for computing GA4GH digests:

<a id="refget.utils.canonical_str"></a>

### `canonical_str`

```python
def canonical_str(item: dict) -> bytes
```

Convert a dict into a canonical string representation
