Skip to content

Data Models

The refget package uses Pydantic and SQLModel for data validation and database ORM. These models represent the core data structures for sequence collections, DRS objects, and related metadata.

DrsObject (base)
└── FastaDrsObject (table)
SQLModel (base)
├── SequenceCollection (table)
├── Pangenome (table)
├── Sequence (table)
├── AccessMethod
├── AccessURL
└── Checksum

The primary model representing a GA4GH sequence collection.

class SequenceCollection

Bases: SQLModel

A SQLModel/pydantic model that represents a refget sequence collection.

  • digest (str): Top-level digest of the SequenceCollection.
  • human_readable_names (List[HumanReadableNames]) = Relationship(back_populates='collection')
  • lengths (LengthsAttr) = Relationship(back_populates='collection'): Array of sequence lengths.
  • lengths_digest (str)
  • name_length_pairs (NameLengthPairsAttr) = Relationship(back_populates='collection'): Array of name-length pairs, representing the coordinate system of the collection.
  • name_length_pairs_digest (str)
  • names (NamesAttr) = Relationship(back_populates='collection'): Array of sequence names.
  • names_digest (str)
  • pangenomes (List[Pangenome]) = Relationship(back_populates='collections', link_model=PangenomeCollectionLink)
  • sequences (SequencesAttr) = Relationship(back_populates='collection'): Array of sequence digests.
  • sequences_digest (str)
  • sorted_name_length_pairs_digest (str): Digest of the sorted name-length pairs, representing a unique digest of sort-invariant coordinate system.
  • sorted_sequences (SortedSequencesAttr) = Relationship(back_populates='collection'): Array of sorted sequence digests.
  • sorted_sequences_digest (str)

@classmethod
def from_PySequenceCollection(gtars_seq_col: gtarsSequenceCollection) -> SequenceCollection

Given a PySequenceCollection object (from Rust bindings), create a SequenceCollection object.

Parameters:

  • gtars_seq_col (PySequenceCollection): PySequenceCollection object from Rust bindings.

Returns:

Raises:

  • ImportError: If gtars is not installed (required for this conversion)

@classmethod
def from_dict(seqcol_dict: dict, inherent_attrs: Optional[list] = DEFAULT_INHERENT_ATTRS) -> SequenceCollection

Given a dict representation of a sequence collection, create a SequenceCollection object. This is the primary way to create a SequenceCollection object.

Parameters:

  • seqcol_dict (dict): Dictionary representation of a canonical sequence collection object
  • inherent_attrs (list): List of inherent attributes to digest (default: DEFAULT_INHERENT_ATTRS)

Returns:

@classmethod
def from_fasta_file(fasta_file: str) -> SequenceCollection

Given a FASTA file, create a SequenceCollection object.

Parameters:

  • fasta_file (str): Path to a FASTA file

Returns:

Raises:

  • ImportError: If gtars is not installed (required for FASTA processing)

def itemwise(limit=None)

Converts object into a list of dictionaries, one for each sequence in the collection.

def level1()

Converts object into dict of level 1 representation of the SequenceCollection.

Returns attribute digests for most attributes, but returns raw values for passthru attributes. Note: Passthru handling for dict-based construction happens in seqcol_dict_to_level1_dict(). When passthru attributes are added to the database model, return .value instead of .digest here.

def level2()

Converts object into dict of level 2 representation of the SequenceCollection.

A DRS object specialized for FASTA files, storing file metadata and FAI index information.

class FastaDrsObject

Bases: DrsObject

A DRS object specialized for FASTA sequence files. Stores file metadata including size, checksums (SHA-256, MD5, and refget sequence collection digest), and creation time. The refget digest serves as the object ID, enabling content-addressable retrieval.

  • extra_line_bytes (Optional[int]) = None
  • id (str)
  • line_bases (Optional[int]) = None
  • offsets (Optional[List[int]]) = None
  • self_uri (Optional[str]) = None

@classmethod
def from_fasta_file(fasta_file: str, digest: str = None) -> FastaDrsObject

Given a FASTA file, create a FastaDrsObject object, return a populated FastaDrsObject with computed size and checksum.

Parameters:

  • fasta_file (str): Path to a FASTA file
  • digest (str): The refget digest of the sequence collection (optional). If not included, it will be computed (default: None)

Returns:

Raises:

  • ImportError: If gtars is not installed (required for FASTA processing)

def to_response(base_uri: str = None) -> FastaDrsObject

Return a copy of this object with self_uri populated for API response.

Parameters:

  • base_uri (str): Base URI for the DRS service (e.g., "drs://seqcolapi.databio.org") If not provided, returns self unchanged. (default: None)

Returns:

Base model for GA4GH Data Repository Service (DRS) objects.

class DrsObject

Bases: SQLModel

A data object representing a single blob of bytes with metadata, checksums, and access methods. DRS objects are self-contained and provide all information needed for clients to retrieve the data. Conforms to GA4GH Data Repository Service (DRS) specification v1.4.0.

  • access_methods (List[AccessMethod])
  • aliases (List[str])
  • checksums (List[Checksum])
  • created_time (datetime)
  • description (Optional[str]) = None
  • id (str)
  • mime_type (Optional[str]) = None
  • name (Optional[str]) = None
  • self_uri (str)
  • size (int)
  • updated_time (Optional[datetime]) = None
  • version (Optional[str]) = None

@classmethod
def coerce_access_methods(v)

Coerce dicts to AccessMethod objects when loading from JSON.

@classmethod
def coerce_checksums(v)

Coerce dicts to Checksum objects when loading from JSON.

def serialize_access_methods(v)

Serialize AccessMethod objects (or dicts) to dicts for JSON output.

def serialize_checksums(v)

Serialize Checksum objects (or dicts) to dicts for JSON output.

A collection of sequence collections representing a pangenome.

class Pangenome

Bases: SQLModel

  • collections (List[SequenceCollection]) = Relationship(back_populates='pangenomes', link_model=PangenomeCollectionLink)
  • collections_digest (str)
  • digest (str)
  • names (CollectionNamesAttr) = Relationship(back_populates='pangenome')
  • names_digest (str)

@classmethod
def from_dict(pangenome_obj: dict, inherent_attrs: Optional[list] = None) -> Pangenome

Given a dict representation of a pangenome, create a Pangenome object. This is the primary way to create a Pangenome object.

Parameters:

  • pangenome_obj (dict): Dictionary representation of a canonical pangenome object

Returns:

def level1()

Converts object into dict of level 1 representation of the Pangenome.

def level2()

Converts object into dict of level 2 representation of the Pangenome.

def level3()

Converts object into dict of level 3 representation of the Pangenome.

def level4()

Converts object into dict of level 4 representation of the Pangenome.

An individual sequence with its digest and content.

class Sequence

Bases: SQLModel

  • digest (str)
  • length (int)
  • sequence (str)

Describes how to access object bytes (protocol type, URL, region).

class AccessMethod

Bases: SQLModel

Describes a method for accessing object bytes, including the protocol type (e.g., https, s3, gs) and either a direct URL or an access_id for the /access endpoint. At least one of access_url or access_id must be provided.

DRS 1.5.0 adds the 'cloud' field to explicitly specify the cloud provider.

  • access_id (Optional[str]) = None
  • access_url (Optional[AccessURL]) = None
  • cloud (Optional[str]) = None
  • region (Optional[str]) = None
  • type (Literal['s3', 'gs', 'ftp', 'gsiftp', 'globus', 'htsget', 'https', 'file'])

A fully resolvable URL with optional headers for authentication.

class AccessURL

Bases: SQLModel

A fully resolvable URL that can be used to fetch the actual object bytes. Optionally includes headers (e.g., authorization tokens) required for access.

  • headers (Optional[List[str]]) = None
  • url (str)

A checksum for data integrity verification.

class Checksum

Bases: SQLModel

A checksum for data integrity verification. The type field indicates the hash algorithm (e.g., "sha-256", "md5") and the checksum field contains the hex-string encoded hash value.

  • checksum (str)
  • type (str)

Pagination metadata for list endpoints.

class PaginationResult

Bases: BaseModel

  • page (int) = 0
  • page_size (int) = 10
  • total (int)

Paginated sequence collection results.

class ResultsSequenceCollections

Bases: BaseModel

Sequence collection results with pagination

Results from Jaccard similarity calculations.

class Similarities

Bases: BaseModel

Model to contain results from similarities calculations

  • pagination (PaginationResult)
  • reference_digest (Optional[str]) = None
  • similarities (List[Dict[str, Any]])

These models store individual attributes of sequence collections in normalized database tables:

class NamesAttr

Bases: SQLModel

  • collection (List[SequenceCollection]) = Relationship(back_populates='names')
  • digest (str)
  • value (list)

class LengthsAttr

Bases: SQLModel

  • collection (List[SequenceCollection]) = Relationship(back_populates='lengths')
  • digest (str)
  • value (list)

class SequencesAttr

Bases: SQLModel

  • collection (List[SequenceCollection]) = Relationship(back_populates='sequences')
  • digest (str)
  • value (list)

class NameLengthPairsAttr

Bases: SQLModel

  • collection (List[SequenceCollection]) = Relationship(back_populates='name_length_pairs')
  • digest (str)
  • value (list)

Creating a SequenceCollection from a FASTA file

Section titled “Creating a SequenceCollection from a FASTA file”
from refget.models import SequenceCollection
# From a FASTA file (requires gtars)
seqcol = SequenceCollection.from_fasta_file("genome.fa")
# Access different representations
print(seqcol.digest) # Top-level digest
print(seqcol.level1()) # Attribute digests
print(seqcol.level2()) # Full arrays
print(seqcol.itemwise()) # Per-sequence dicts

Creating a SequenceCollection from a dictionary

Section titled “Creating a SequenceCollection from a dictionary”
from refget.models import SequenceCollection
seqcol_dict = {
"names": ["chr1", "chr2"],
"lengths": [1000, 2000],
"sequences": ["SQ.abc123...", "SQ.def456..."]
}
seqcol = SequenceCollection.from_dict(seqcol_dict)
from refget.models import FastaDrsObject
# From a FASTA file
drs_obj = FastaDrsObject.from_fasta_file("genome.fa")
# Access DRS metadata
print(drs_obj.id) # Sequence collection digest
print(drs_obj.size) # File size in bytes
print(drs_obj.checksums) # SHA-256, MD5
print(drs_obj.access_methods) # How to download