The `refget` package uses Pydantic and SQLModel for data validation and database ORM. These models represent the core data structures for sequence collections, DRS objects, and related metadata.

:::tip[Data models]
Data Models are only needed if you want to develop new packages that rely on the refget Python API.

:::


## Model hierarchy

```
DrsObject (base)
└── FastaDrsObject (table)

SQLModel (base)
├── SequenceCollection (table)
├── Pangenome (table)
├── Sequence (table)
├── AccessMethod
├── AccessURL
└── Checksum
```

## Core Models

### SequenceCollection

The primary model representing a GA4GH sequence collection.

<a id="refget.models.SequenceCollection"></a>

#### `SequenceCollection`

```python
class SequenceCollection
```

Bases: `SQLModel`

A SQLModel/pydantic model that represents a refget sequence collection.

##### Attributes

- <a id="refget.models.SequenceCollection.digest"></a>**`digest`** (*str*): Top-level digest of the SequenceCollection.
- <a id="refget.models.SequenceCollection.human_readable_names"></a>**`human_readable_names`** (*List[HumanReadableNames]*) = `Relationship(back_populates='collection')`
- <a id="refget.models.SequenceCollection.lengths"></a>**`lengths`** (*[LengthsAttr](#refget.models.LengthsAttr)*) = `Relationship(back_populates='collection')`: Array of sequence lengths.
- <a id="refget.models.SequenceCollection.lengths_digest"></a>**`lengths_digest`** (*str*)
- <a id="refget.models.SequenceCollection.name_length_pairs"></a>**`name_length_pairs`** (*[NameLengthPairsAttr](#refget.models.NameLengthPairsAttr)*) = `Relationship(back_populates='collection')`: Array of name-length pairs, representing the coordinate system of the collection.
- <a id="refget.models.SequenceCollection.name_length_pairs_digest"></a>**`name_length_pairs_digest`** (*str*)
- <a id="refget.models.SequenceCollection.names"></a>**`names`** (*[NamesAttr](#refget.models.NamesAttr)*) = `Relationship(back_populates='collection')`: Array of sequence names.
- <a id="refget.models.SequenceCollection.names_digest"></a>**`names_digest`** (*str*)
- <a id="refget.models.SequenceCollection.pangenomes"></a>**`pangenomes`** (*List[[Pangenome](#refget.models.Pangenome)]*) = `Relationship(back_populates='collections', link_model=PangenomeCollectionLink)`
- <a id="refget.models.SequenceCollection.sequences"></a>**`sequences`** (*[SequencesAttr](#refget.models.SequencesAttr)*) = `Relationship(back_populates='collection')`: Array of sequence digests.
- <a id="refget.models.SequenceCollection.sequences_digest"></a>**`sequences_digest`** (*str*)
- <a id="refget.models.SequenceCollection.sorted_name_length_pairs_digest"></a>**`sorted_name_length_pairs_digest`** (*str*): Digest of the sorted name-length pairs, representing a unique digest of sort-invariant coordinate system.
- <a id="refget.models.SequenceCollection.sorted_sequences"></a>**`sorted_sequences`** (*SortedSequencesAttr*) = `Relationship(back_populates='collection')`: Array of sorted sequence digests.
- <a id="refget.models.SequenceCollection.sorted_sequences_digest"></a>**`sorted_sequences_digest`** (*str*)

##### Class methods

<a id="refget.models.SequenceCollection.from_PySequenceCollection"></a>

###### `from_PySequenceCollection`

```python
@classmethod
def from_PySequenceCollection(gtars_seq_col: gtarsSequenceCollection) -> SequenceCollection
```

Given a PySequenceCollection object (from Rust bindings), create a SequenceCollection object.

**Parameters:**

- `gtars_seq_col` (*PySequenceCollection*): PySequenceCollection object from Rust bindings.

**Returns:**

- *[SequenceCollection](#refget.models.SequenceCollection)*: The SequenceCollection object.

**Raises:**

- *ImportError*: If gtars is not installed (required for this conversion)

<a id="refget.models.SequenceCollection.from_dict"></a>

###### `from_dict`

```python
@classmethod
def from_dict(seqcol_dict: dict, inherent_attrs: Optional[list] = DEFAULT_INHERENT_ATTRS) -> SequenceCollection
```

Given a dict representation of a sequence collection, create a SequenceCollection object.
This is the primary way to create a SequenceCollection object.

**Parameters:**

- `seqcol_dict` (*dict*): Dictionary representation of a canonical sequence collection object
- `inherent_attrs` (*list*): List of inherent attributes to digest (default: `DEFAULT_INHERENT_ATTRS`)

**Returns:**

- *[SequenceCollection](#refget.models.SequenceCollection)*: The SequenceCollection object

<a id="refget.models.SequenceCollection.from_fasta_file"></a>

###### `from_fasta_file`

```python
@classmethod
def from_fasta_file(fasta_file: str) -> SequenceCollection
```

Given a FASTA file, create a SequenceCollection object.

**Parameters:**

- `fasta_file` (*str*): Path to a FASTA file

**Returns:**

- *[SequenceCollection](#refget.models.SequenceCollection)*: The SequenceCollection object

**Raises:**

- *ImportError*: If gtars is not installed (required for FASTA processing)

##### Methods

<a id="refget.models.SequenceCollection.itemwise"></a>

###### `itemwise`

```python
def itemwise(limit=None)
```

Converts object into a list of dictionaries, one for each sequence in the collection.

<a id="refget.models.SequenceCollection.level1"></a>

###### `level1`

```python
def level1()
```

Converts object into dict of level 1 representation of the SequenceCollection.

Returns attribute digests for most attributes, but returns raw values for passthru attributes.
Note: Passthru handling for dict-based construction happens in seqcol_dict_to_level1_dict().
When passthru attributes are added to the database model, return .value instead of .digest here.

<a id="refget.models.SequenceCollection.level2"></a>

###### `level2`

```python
def level2()
```

Converts object into dict of level 2 representation of the SequenceCollection.

### FastaDrsObject

A DRS object specialized for FASTA files, storing file metadata and FAI index information.

<a id="refget.models.FastaDrsObject"></a>

#### `FastaDrsObject`

```python
class FastaDrsObject
```

Bases: [`DrsObject`](#refget.models.DrsObject)

A DRS object specialized for FASTA sequence files. Stores file metadata including
size, checksums (SHA-256, MD5, and refget sequence collection digest), and creation time.
The refget digest serves as the object ID, enabling content-addressable retrieval.

##### Attributes

- <a id="refget.models.FastaDrsObject.extra_line_bytes"></a>**`extra_line_bytes`** (*Optional[int]*) = `None`
- <a id="refget.models.FastaDrsObject.id"></a>**`id`** (*str*)
- <a id="refget.models.FastaDrsObject.line_bases"></a>**`line_bases`** (*Optional[int]*) = `None`
- <a id="refget.models.FastaDrsObject.offsets"></a>**`offsets`** (*Optional[List[int]]*) = `None`
- <a id="refget.models.FastaDrsObject.self_uri"></a>**`self_uri`** (*Optional[str]*) = `None`

##### Class methods

<a id="refget.models.FastaDrsObject.from_fasta_file"></a>

###### `from_fasta_file`

```python
@classmethod
def from_fasta_file(fasta_file: str, digest: str = None) -> FastaDrsObject
```

Given a FASTA file, create a FastaDrsObject object,
return a populated FastaDrsObject with computed size and checksum.

**Parameters:**

- `fasta_file` (*str*): Path to a FASTA file
- `digest` (*str*): The refget digest of the sequence collection
  (optional). If not included, it will be computed (default: `None`)

**Returns:**

- *[FastaDrsObject](#refget.models.FastaDrsObject)*: The FastaDrsObject object

**Raises:**

- *ImportError*: If gtars is not installed (required for FASTA processing)

##### Methods

<a id="refget.models.FastaDrsObject.to_response"></a>

###### `to_response`

```python
def to_response(base_uri: str = None) -> FastaDrsObject
```

Return a copy of this object with self_uri populated for API response.

**Parameters:**

- `base_uri` (*str*): Base URI for the DRS service (e.g., "drs://seqcolapi.databio.org")
       If not provided, returns self unchanged. (default: `None`)

**Returns:**

- *[FastaDrsObject](#refget.models.FastaDrsObject)*: FastaDrsObject with self_uri populated

### DrsObject

Base model for GA4GH Data Repository Service (DRS) objects.

<a id="refget.models.DrsObject"></a>

#### `DrsObject`

```python
class DrsObject
```

Bases: `SQLModel`

A data object representing a single blob of bytes with metadata, checksums, and access methods.
DRS objects are self-contained and provide all information needed for clients to retrieve the data.
Conforms to GA4GH Data Repository Service (DRS) specification v1.4.0.

##### Attributes

- <a id="refget.models.DrsObject.access_methods"></a>**`access_methods`** (*List[[AccessMethod](#refget.models.AccessMethod)]*)
- <a id="refget.models.DrsObject.aliases"></a>**`aliases`** (*List[str]*)
- <a id="refget.models.DrsObject.checksums"></a>**`checksums`** (*List[[Checksum](#refget.models.Checksum)]*)
- <a id="refget.models.DrsObject.created_time"></a>**`created_time`** (*datetime*)
- <a id="refget.models.DrsObject.description"></a>**`description`** (*Optional[str]*) = `None`
- <a id="refget.models.DrsObject.id"></a>**`id`** (*str*)
- <a id="refget.models.DrsObject.mime_type"></a>**`mime_type`** (*Optional[str]*) = `None`
- <a id="refget.models.DrsObject.name"></a>**`name`** (*Optional[str]*) = `None`
- <a id="refget.models.DrsObject.self_uri"></a>**`self_uri`** (*str*)
- <a id="refget.models.DrsObject.size"></a>**`size`** (*int*)
- <a id="refget.models.DrsObject.updated_time"></a>**`updated_time`** (*Optional[datetime]*) = `None`
- <a id="refget.models.DrsObject.version"></a>**`version`** (*Optional[str]*) = `None`

##### Class methods

<a id="refget.models.DrsObject.coerce_access_methods"></a>

###### `coerce_access_methods`

```python
@classmethod
def coerce_access_methods(v)
```

Coerce dicts to AccessMethod objects when loading from JSON.

<a id="refget.models.DrsObject.coerce_checksums"></a>

###### `coerce_checksums`

```python
@classmethod
def coerce_checksums(v)
```

Coerce dicts to Checksum objects when loading from JSON.

##### Methods

<a id="refget.models.DrsObject.serialize_access_methods"></a>

###### `serialize_access_methods`

```python
def serialize_access_methods(v)
```

Serialize AccessMethod objects (or dicts) to dicts for JSON output.

<a id="refget.models.DrsObject.serialize_checksums"></a>

###### `serialize_checksums`

```python
def serialize_checksums(v)
```

Serialize Checksum objects (or dicts) to dicts for JSON output.

### Pangenome

A collection of sequence collections representing a pangenome.

<a id="refget.models.Pangenome"></a>

#### `Pangenome`

```python
class Pangenome
```

Bases: `SQLModel`

##### Attributes

- <a id="refget.models.Pangenome.collections"></a>**`collections`** (*List[[SequenceCollection](#refget.models.SequenceCollection)]*) = `Relationship(back_populates='pangenomes', link_model=PangenomeCollectionLink)`
- <a id="refget.models.Pangenome.collections_digest"></a>**`collections_digest`** (*str*)
- <a id="refget.models.Pangenome.digest"></a>**`digest`** (*str*)
- <a id="refget.models.Pangenome.names"></a>**`names`** (*CollectionNamesAttr*) = `Relationship(back_populates='pangenome')`
- <a id="refget.models.Pangenome.names_digest"></a>**`names_digest`** (*str*)

##### Class methods

<a id="refget.models.Pangenome.from_dict"></a>

###### `from_dict`

```python
@classmethod
def from_dict(pangenome_obj: dict, inherent_attrs: Optional[list] = None) -> Pangenome
```

Given a dict representation of a pangenome, create a Pangenome object.
This is the primary way to create a Pangenome object.

**Parameters:**

- `pangenome_obj` (*dict*): Dictionary representation of a canonical pangenome object

**Returns:**

- *[Pangenome](#refget.models.Pangenome)*: The Pangenome object

##### Methods

<a id="refget.models.Pangenome.level1"></a>

###### `level1`

```python
def level1()
```

Converts object into dict of level 1 representation of the Pangenome.

<a id="refget.models.Pangenome.level2"></a>

###### `level2`

```python
def level2()
```

Converts object into dict of level 2 representation of the Pangenome.

<a id="refget.models.Pangenome.level3"></a>

###### `level3`

```python
def level3()
```

Converts object into dict of level 3 representation of the Pangenome.

<a id="refget.models.Pangenome.level4"></a>

###### `level4`

```python
def level4()
```

Converts object into dict of level 4 representation of the Pangenome.

### Sequence

An individual sequence with its digest and content.

<a id="refget.models.Sequence"></a>

#### `Sequence`

```python
class Sequence
```

Bases: `SQLModel`

##### Attributes

- <a id="refget.models.Sequence.digest"></a>**`digest`** (*str*)
- <a id="refget.models.Sequence.length"></a>**`length`** (*int*)
- <a id="refget.models.Sequence.sequence"></a>**`sequence`** (*str*)

## Supporting Models

### AccessMethod

Describes how to access object bytes (protocol type, URL, region).

<a id="refget.models.AccessMethod"></a>

#### `AccessMethod`

```python
class AccessMethod
```

Bases: `SQLModel`

Describes a method for accessing object bytes, including the protocol type
(e.g., https, s3, gs) and either a direct URL or an access_id for the /access endpoint.
At least one of access_url or access_id must be provided.

DRS 1.5.0 adds the 'cloud' field to explicitly specify the cloud provider.

##### Attributes

- <a id="refget.models.AccessMethod.access_id"></a>**`access_id`** (*Optional[str]*) = `None`
- <a id="refget.models.AccessMethod.access_url"></a>**`access_url`** (*Optional[[AccessURL](#refget.models.AccessURL)]*) = `None`
- <a id="refget.models.AccessMethod.cloud"></a>**`cloud`** (*Optional[str]*) = `None`
- <a id="refget.models.AccessMethod.region"></a>**`region`** (*Optional[str]*) = `None`
- <a id="refget.models.AccessMethod.type"></a>**`type`** (*Literal['s3', 'gs', 'ftp', 'gsiftp', 'globus', 'htsget', 'https', 'file']*)

### AccessURL

A fully resolvable URL with optional headers for authentication.

<a id="refget.models.AccessURL"></a>

#### `AccessURL`

```python
class AccessURL
```

Bases: `SQLModel`

A fully resolvable URL that can be used to fetch the actual object bytes.
Optionally includes headers (e.g., authorization tokens) required for access.

##### Attributes

- <a id="refget.models.AccessURL.headers"></a>**`headers`** (*Optional[List[str]]*) = `None`
- <a id="refget.models.AccessURL.url"></a>**`url`** (*str*)

### Checksum

A checksum for data integrity verification.

<a id="refget.models.Checksum"></a>

#### `Checksum`

```python
class Checksum
```

Bases: `SQLModel`

A checksum for data integrity verification. The type field indicates the hash algorithm
(e.g., "sha-256", "md5") and the checksum field contains the hex-string encoded hash value.

##### Attributes

- <a id="refget.models.Checksum.checksum"></a>**`checksum`** (*str*)
- <a id="refget.models.Checksum.type"></a>**`type`** (*str*)

## Response Models

### PaginationResult

Pagination metadata for list endpoints.

<a id="refget.models.PaginationResult"></a>

#### `PaginationResult`

```python
class PaginationResult
```

Bases: `BaseModel`

##### Attributes

- <a id="refget.models.PaginationResult.page"></a>**`page`** (*int*) = `0`
- <a id="refget.models.PaginationResult.page_size"></a>**`page_size`** (*int*) = `10`
- <a id="refget.models.PaginationResult.total"></a>**`total`** (*int*)

### ResultsSequenceCollections

Paginated sequence collection results.

<a id="refget.models.ResultsSequenceCollections"></a>

#### `ResultsSequenceCollections`

```python
class ResultsSequenceCollections
```

Bases: `BaseModel`

Sequence collection results with pagination

##### Attributes

- <a id="refget.models.ResultsSequenceCollections.pagination"></a>**`pagination`** (*[PaginationResult](#refget.models.PaginationResult)*)
- <a id="refget.models.ResultsSequenceCollections.results"></a>**`results`** (*Dict[str, dict]*)

### Similarities

Results from Jaccard similarity calculations.

<a id="refget.models.Similarities"></a>

#### `Similarities`

```python
class Similarities
```

Bases: `BaseModel`

Model to contain results from similarities calculations

##### Attributes

- <a id="refget.models.Similarities.pagination"></a>**`pagination`** (*[PaginationResult](#refget.models.PaginationResult)*)
- <a id="refget.models.Similarities.reference_digest"></a>**`reference_digest`** (*Optional[str]*) = `None`
- <a id="refget.models.Similarities.similarities"></a>**`similarities`** (*List[Dict[str, Any]]*)

## Attribute Tables

These models store individual attributes of sequence collections in normalized database tables:

### NamesAttr

<a id="refget.models.NamesAttr"></a>

#### `NamesAttr`

```python
class NamesAttr
```

Bases: `SQLModel`

##### Attributes

- <a id="refget.models.NamesAttr.collection"></a>**`collection`** (*List[[SequenceCollection](#refget.models.SequenceCollection)]*) = `Relationship(back_populates='names')`
- <a id="refget.models.NamesAttr.digest"></a>**`digest`** (*str*)
- <a id="refget.models.NamesAttr.value"></a>**`value`** (*list*)

### LengthsAttr

<a id="refget.models.LengthsAttr"></a>

#### `LengthsAttr`

```python
class LengthsAttr
```

Bases: `SQLModel`

##### Attributes

- <a id="refget.models.LengthsAttr.collection"></a>**`collection`** (*List[[SequenceCollection](#refget.models.SequenceCollection)]*) = `Relationship(back_populates='lengths')`
- <a id="refget.models.LengthsAttr.digest"></a>**`digest`** (*str*)
- <a id="refget.models.LengthsAttr.value"></a>**`value`** (*list*)

### SequencesAttr

<a id="refget.models.SequencesAttr"></a>

#### `SequencesAttr`

```python
class SequencesAttr
```

Bases: `SQLModel`

##### Attributes

- <a id="refget.models.SequencesAttr.collection"></a>**`collection`** (*List[[SequenceCollection](#refget.models.SequenceCollection)]*) = `Relationship(back_populates='sequences')`
- <a id="refget.models.SequencesAttr.digest"></a>**`digest`** (*str*)
- <a id="refget.models.SequencesAttr.value"></a>**`value`** (*list*)

### NameLengthPairsAttr

<a id="refget.models.NameLengthPairsAttr"></a>

#### `NameLengthPairsAttr`

```python
class NameLengthPairsAttr
```

Bases: `SQLModel`

##### Attributes

- <a id="refget.models.NameLengthPairsAttr.collection"></a>**`collection`** (*List[[SequenceCollection](#refget.models.SequenceCollection)]*) = `Relationship(back_populates='name_length_pairs')`
- <a id="refget.models.NameLengthPairsAttr.digest"></a>**`digest`** (*str*)
- <a id="refget.models.NameLengthPairsAttr.value"></a>**`value`** (*list*)

## Usage Examples

### Creating a SequenceCollection from a FASTA file

```python
from refget.models import SequenceCollection

# From a FASTA file (requires gtars)
seqcol = SequenceCollection.from_fasta_file("genome.fa")

# Access different representations
print(seqcol.digest)  # Top-level digest
print(seqcol.level1())  # Attribute digests
print(seqcol.level2())  # Full arrays
print(seqcol.itemwise())  # Per-sequence dicts
```

### Creating a SequenceCollection from a dictionary

```python
from refget.models import SequenceCollection

seqcol_dict = {
    "names": ["chr1", "chr2"],
    "lengths": [1000, 2000],
    "sequences": ["SQ.abc123...", "SQ.def456..."]
}

seqcol = SequenceCollection.from_dict(seqcol_dict)
```

### Creating a FastaDrsObject

```python
from refget.models import FastaDrsObject

# From a FASTA file
drs_obj = FastaDrsObject.from_fasta_file("genome.fa")

# Access DRS metadata
print(drs_obj.id)  # Sequence collection digest
print(drs_obj.size)  # File size in bytes
print(drs_obj.checksums)  # SHA-256, MD5
print(drs_obj.access_methods)  # How to download
```
