Refget Python API Documentation
FASTA Processing
Section titled “FASTA Processing”fasta_to_seqcol_dict
Section titled “fasta_to_seqcol_dict”def fasta_to_seqcol_dict(fasta_file_path: Union[str, Path]) -> dictConvert a FASTA file into a Sequence Collection dict.
Parameters:
fasta_file_path(Union[str, Path]): Path to the FASTA file
Returns:
- dict: A canonical sequence collection dictionary
Raises:
- ImportError: If gtars is not installed (required for FASTA processing)
compare_seqcols
Section titled “compare_seqcols”def compare_seqcols(A: SeqColDict, B: SeqColDict) -> dictWorkhorse comparison function
Parameters:
A(SeqColDict): Sequence collection AB(SeqColDict): Sequence collection B
Returns:
- dict: Following formal seqcol specification comparison function return value
calc_jaccard_similarities
Section titled “calc_jaccard_similarities”def calc_jaccard_similarities(A: SeqColDict, B: SeqColDict) -> dict[str, float]Takes two sequence collections and calculates jaccard similarties for all attributes
Parameters:
A(SeqColDict): Sequence collection AB(SeqColDict): Sequence collection B
Returns:
dict(dict[str, float]): Jaccard similarities for all attributes
validate_seqcol
Section titled “validate_seqcol”def validate_seqcol(seqcol_obj: SeqColDict, schema=None) -> boolValidate a seqcol object against the seqcol schema. Returns True if valid, raises InvalidSeqColError if not, which enumerates the errors. Retrieve individual errors with exception.errors
validate_seqcol_bool
Section titled “validate_seqcol_bool”def validate_seqcol_bool(seqcol_obj: SeqColDict, schema=None) -> boolValidate a seqcol object against the seqcol schema. Returns True if valid, False if not.
To enumerate the errors, use validate_seqcol instead.
FastAPI Integration
Section titled “FastAPI Integration”create_refget_router
Section titled “create_refget_router”def create_refget_router(sequences: bool = False, collections: bool = True, pangenomes: bool = False, fasta_drs: bool = False, compliance: bool = True, refget_store_url: str = None, mount_prefix: str = '') -> APIRouterCreate a FastAPI router for the sequence collection API. This router provides endpoints for retrieving and comparing sequence collections. You can choose which endpoints to include by setting the sequences, collections, pangenomes, or fasta_drs flags.
Parameters:
sequences(bool): Include sequence endpoints (default:False)collections(bool): Include sequence collection endpoints (default:True)pangenomes(bool): Include pangenome endpoints (default:False)fasta_drs(bool): Include FASTA DRS endpoints (default:False)refget_store_url(str): URL of backing RefgetStore (e.g., s3://bucket/store/) (default:None)mount_prefix(str): The path prefix this router will be included under, when it is included withinclude_router(..., prefix=...)rather than mounted as a sub-application. Only used so the compliance endpoints self-target the seqcol service instead of the server root. Leave empty when mounting an app (scope["root_path"]covers it). (default:'')
Returns:
- APIRouter: A FastAPI router with the specified endpoints
Examples:
app.include_router(create_refget_router(fasta_drs=True), prefix="/seqcol")Client Classes
Section titled “Client Classes”The client module provides interfaces for interacting with refget-compliant servers.
SequenceClient
Section titled “SequenceClient”class SequenceClient(urls: list[str] = ['https://www.ebi.ac.uk/ena/cram'], raise_errors: Optional[bool] = None)Bases: RefgetClient
A client for interacting with a refget sequences API.
Initializes the sequences client.
Parameters:
urls(list): A list of base URLs of the sequences API. Defaults to ["https://www.ebi.ac.uk/ena/cram/sequence/"]. (default:['https://www.ebi.ac.uk/ena/cram'])raise_errors(bool): Whether to raise errors or log them. Defaults to None, which will guess. (default:None)
Attributes: urls (list): The list of base URLs of the sequences API.
Attributes
Section titled “Attributes”Methods
Section titled “Methods”get_metadata
Section titled “get_metadata”def get_metadata(digest: str) -> Optional[dict]Retrieves metadata for a given sequence digest.
Parameters:
digest(str): The digest of the sequence.
Returns:
- dict: The metadata.
get_sequence
Section titled “get_sequence”def get_sequence(digest: str, start: Optional[int] = None, end: Optional[int] = None) -> Optional[str]Retrieves a sequence for a given digest.
Parameters:
digest(str): The digest of the sequence.
Returns:
- str: The sequence.
SequenceCollectionClient
Section titled “SequenceCollectionClient”class SequenceCollectionClient(urls: list[str] = ['https://seqcolapi.databio.org'], raise_errors: Optional[bool] = None)Bases: RefgetClient
A client for interacting with a refget sequence collections API.
Initializes the sequence collection client.
Parameters:
urls(list): A list of base URLs of the sequence collection API. Defaults to ["https://seqcolapi.databio.org"]. (default:['https://seqcolapi.databio.org'])
Attributes:
urls(list): The list of base URLs of the sequence collection API.
Attributes
Section titled “Attributes”Methods
Section titled “Methods”aliases_for
Section titled “aliases_for”def aliases_for(digest: str, kind: str = 'collection') -> Optional[dict]Reverse lookup: list all (namespace, alias) pairs for a digest.
Parameters:
digest(str): The digest to look up.kind(str): "collection" (default) or "sequence". (default:'collection')
Returns:
- dict: {"digest": ..., "aliases": [[namespace, alias], ...]}.
build_chrom_sizes
Section titled “build_chrom_sizes”def build_chrom_sizes(digest: str) -> strBuild a chrom.sizes file content for a sequence collection.
Format per line: NAME\tLENGTH
Parameters:
digest(str): The sequence collection digest
Returns:
- str: String content of the chrom.sizes file
build_fai
Section titled “build_fai”def build_fai(digest: str) -> strBuild a complete .fai index file content for a FASTA.
FAI format per line: NAME\tLENGTH\tOFFSET\tLINEBASES\tLINEWIDTH
Parameters:
digest(str): The sequence collection digest
Returns:
- str: String content of the .fai file
compare
Section titled “compare”def compare(digest1: str, digest2: str) -> Optional[dict]Compares two sequence collections hosted on the server.
Parameters:
digest1(str): The digest of the first sequence collection.digest2(str): The digest of the second sequence collection.
Returns:
- dict: The JSON response containing the comparison of the two sequence collections.
compare_local
Section titled “compare_local”def compare_local(digest: str, local_collection: dict) -> Optional[dict]Compares a server-hosted sequence collection with a local collection.
Parameters:
digest(str): The digest of the server-hosted sequence collection.local_collection(dict): A level 2 sequence collection representation.
Returns:
- dict: The JSON response containing the comparison.
download_fasta
Section titled “download_fasta”def download_fasta(digest: str, dest_path: str = None, access_id: str = None) -> strDownload the FASTA file to a local path.
Parameters:
digest(str): The sequence collection digestdest_path(str): Destination file path. If None, uses object name. (default:None)access_id(str): Specific access method to use. If None, tries all. (default:None)
Returns:
- str: Path to downloaded file
Raises:
- ValueError: If no access methods available or specified access_id not found
download_fasta_to_store
Section titled “download_fasta_to_store”def download_fasta_to_store(digest: str, store: 'RefgetStore', access_id: str = None, temp_dir: str = None, namespaces: Optional[list[str]] = None) -> strDownload the FASTA file and import it into a RefgetStore.
This method downloads the FASTA file from the DRS endpoint and immediately imports it into the provided RefgetStore, enabling local sequence retrieval by digest without re-downloading.
Parameters:
digest(str): The sequence collection digeststore(RefgetStore): The RefgetStore instance to import intoaccess_id(str): Specific access method to use. If None, tries all. (default:None)temp_dir(str): Directory for temporary download. If None, uses system temp. (default:None)
Returns:
- str: The collection digest of the imported sequences
Raises:
- ValueError: If no access methods available or specified access_id not found
- ImportError: If gtars/RefgetStore is not available
get_attribute
Section titled “get_attribute”def get_attribute(attribute: str, digest: str) -> Optional[dict]Retrieves a specific attribute value by its digest.
Parameters:
attribute(str): The attribute name (e.g., "names", "lengths", "sequences").digest(str): The level 1 digest of the attribute.
Returns:
- dict: The JSON response containing the attribute value.
get_collection
Section titled “get_collection”def get_collection(digest: str, level: int = 2) -> Optional[dict]Retrieves a sequence collection for a given digest and detail level.
Parameters:
digest(str): The digest of the sequence collection.level(int): The level of detail for the sequence collection. Defaults to 2. (default:2)
Returns:
- dict: The JSON response containing the sequence collection.
get_fasta
Section titled “get_fasta”def get_fasta(digest: str) -> Optional[dict]Get DRS object metadata for a FASTA file.
Parameters:
digest(str): The sequence collection digest (which is also the DRS object ID)
Returns:
- dict: DRS object with id, self_uri, size, checksums, access_methods, etc.
get_fasta_index
Section titled “get_fasta_index”def get_fasta_index(digest: str) -> Optional[dict]Get FAI index data for a FASTA file.
Parameters:
digest(str): The sequence collection digest
Returns:
- dict: Dict with line_bases, extra_line_bytes, offsets
get_fhr
Section titled “get_fhr”def get_fhr(digest: str) -> Optional[dict]Get FHR metadata for a collection.
Parameters:
digest(str): The collection digest.
Returns:
- dict: FHR metadata, or None if not found.
get_refget_store
Section titled “get_refget_store”def get_refget_store(cache_dir: str) -> 'RefgetStore'Get a RefgetStore instance connected to the server's backing store.
Parameters:
cache_dir(str): Local directory for caching store data
Returns:
- RefgetStore: RefgetStore instance loaded from remote
Raises:
- ValueError: If server doesn't have a RefgetStore configured
- ImportError: If gtars is not installed
get_refget_store_url
Section titled “get_refget_store_url”def get_refget_store_url() -> Optional[str]Discover RefgetStore URL from service-info if available.
Returns:
- str: The RefgetStore URL if configured, None otherwise.
get_regions
Section titled “get_regions”def get_regions(digest: str, regions: list) -> Optional[list]Extract region substrings from a server-hosted sequence collection.
Posts a list of regions to the server's region-extraction endpoint and returns the structured results. Requires a store-backed server (the database backend responds with HTTP 501).
Parameters:
digest(str): The collection digest to extract regions from.regions(list): A list of {"chrom", "start", "end"} dicts.
Returns:
- list: A list of {"chrom_name", "start", "end", "sequence"} dicts.
is_aliases_enabled
Section titled “is_aliases_enabled”def is_aliases_enabled() -> boolCheck if alias endpoints are advertised in service-info.
Returns:
- bool: True if aliases are enabled, False otherwise.
is_fasta_drs_enabled
Section titled “is_fasta_drs_enabled”def is_fasta_drs_enabled() -> boolCheck if FastaDRS endpoints are available.
Returns:
- bool: True if FastaDRS is enabled, False otherwise.
is_fhr_enabled
Section titled “is_fhr_enabled”def is_fhr_enabled() -> boolCheck if FHR metadata is advertised in service-info.
Returns:
- bool: True if FHR metadata is enabled, False otherwise.
list_alias_namespaces
Section titled “list_alias_namespaces”def list_alias_namespaces(kind: str = 'collection') -> Optional[dict]List alias namespaces for the given kind.
Parameters:
kind(str): "collection" (default) or "sequence". (default:'collection')
Returns:
- dict: {"namespaces": [...]}.
list_aliases
Section titled “list_aliases”def list_aliases(namespace: str, kind: str = 'collection') -> Optional[dict]List aliases within a namespace.
Parameters:
namespace(str): The alias namespace.kind(str): "collection" (default) or "sequence". (default:'collection')
Returns:
- dict: {"namespace": ..., "aliases": [...]}.
list_attributes
Section titled “list_attributes”def list_attributes(attribute: str, page: Optional[int] = None, page_size: Optional[int] = None) -> Optional[dict]Lists all available values for a given attribute with optional paging support.
Parameters:
attribute(str): The attribute to list values for.page(int): The page number to retrieve. Defaults to None. (default:None)page_size(int): The number of items per page. Defaults to None. (default:None)
Returns:
- dict: The JSON response containing the list of available values for the attribute.
list_collections
Section titled “list_collections”def list_collections(page: Optional[int] = None, page_size: Optional[int] = None, **filters) -> Optional[dict]Lists all available sequence collections with optional paging and attribute filtering support.
Parameters:
page(int): The page number to retrieve. Defaults to None. (default:None)page_size(int): The number of items per page. Defaults to None. (default:None)**filters(Any): Optional attribute filters (e.g., names="abc123", lengths="def456"). Values should be level 1 digests of the attributes. (default:{})
Returns:
- dict: The JSON response containing the list of available sequence collections.
list_fhr
Section titled “list_fhr”def list_fhr() -> Optional[dict]List collections that have FHR metadata.
Returns:
- dict: {"collections": [...]}.
resolve_alias
Section titled “resolve_alias”def resolve_alias(namespace: str, alias: str, kind: str = 'collection') -> Optional[dict]Resolve a namespace:alias to a digest.
Parameters:
namespace(str): The alias namespace.alias(str): The alias name.kind(str): "collection" (default) or "sequence". (default:'collection')
Returns:
- dict: {"namespace": ..., "alias": ..., "digest": ...} or None.
service_info
Section titled “service_info”def service_info() -> Optional[dict]Retrieves information about the service.
Returns:
- dict: The service information.
write_chrom_sizes
Section titled “write_chrom_sizes”def write_chrom_sizes(digest: str, dest_path: str) -> strWrite a chrom.sizes file for a sequence collection.
Parameters:
digest(str): The sequence collection digestdest_path(str): Path to write the chrom.sizes file
Returns:
- str: Path to the written file
write_fai
Section titled “write_fai”def write_fai(digest: str, dest_path: str) -> strWrite a .fai index file for a FASTA.
Parameters:
digest(str): The sequence collection digestdest_path(str): Path to write the .fai file
Returns:
- str: Path to the written file
FastaDrsClient
Section titled “FastaDrsClient”class FastaDrsClient(urls: list[str] = ['https://seqcolapi.databio.org/fasta'], raise_errors: Optional[bool] = None)Bases: RefgetClient
A client for interacting with FASTA files via GA4GH DRS endpoints.
Initializes the FASTA DRS client.
Parameters:
urls(list): A list of base URLs of the FASTA DRS API. Defaults to ["https://seqcolapi.databio.org/fasta"]. (default:['https://seqcolapi.databio.org/fasta'])raise_errors(bool): Whether to raise errors or log them. Defaults to None, which will guess. (default:None)
Attributes:
urls(list): The list of base URLs of the FASTA DRS API.
Attributes
Section titled “Attributes”Methods
Section titled “Methods”build_fai
Section titled “build_fai”def build_fai(digest: str, seqcol_client: 'SequenceCollectionClient' = None) -> strBuild a complete .fai index file content for a FASTA.
FAI format per line: NAME LENGTH OFFSET LINEBASES LINEWIDTH
Parameters:
digest(str): The sequence collection digestseqcol_client(SequenceCollectionClient): SequenceCollectionClient to use. If None, uses parent client or creates one. (default:None)
Returns:
- str: String content of the .fai file
download
Section titled “download”def download(digest: str, dest_path: str = None, access_id: str = None) -> strDownload the FASTA file to a local path.
Parameters:
digest(str): The sequence collection digestdest_path(str): Destination file path. If None, uses object name. (default:None)access_id(str): Specific access method to use. If None, tries all. (default:None)
Returns:
- str: Path to downloaded file
Raises:
- ValueError: If no access methods available or specified access_id not found
download_to_store
Section titled “download_to_store”def download_to_store(digest: str, store: 'RefgetStore', access_id: str = None, temp_dir: str = None, namespaces: Optional[list[str]] = None) -> strDownload the FASTA file and import it into a RefgetStore.
This method downloads the FASTA file from the DRS endpoint and immediately imports it into the provided RefgetStore, enabling local sequence retrieval by digest without re-downloading.
Parameters:
digest(str): The sequence collection digeststore(RefgetStore): The RefgetStore instance to import intoaccess_id(str): Specific access method to use. If None, tries all. (default:None)temp_dir(str): Directory for temporary download. If None, uses system temp. (default:None)namespaces(list[str]): Namespace prefixes to extract aliases from FASTA headers when importing into the store. (default:None)
Returns:
- str: The collection digest of the imported sequences
Raises:
- ValueError: If no access methods available or specified access_id not found
- ImportError: If gtars/RefgetStore is not available
get_access_url
Section titled “get_access_url”def get_access_url(digest: str, access_id: str) -> Optional[dict]Get access URL for a specific access method.
Parameters:
digest(str): The sequence collection digestaccess_id(str): The access ID from the access method
Returns:
- dict: Access URL object
get_index
Section titled “get_index”def get_index(digest: str) -> Optional[dict]Get FAI index data for a FASTA file.
Parameters:
digest(str): The sequence collection digest
Returns:
- dict: Dict with line_bases, extra_line_bytes, offsets
get_object
Section titled “get_object”def get_object(digest: str) -> Optional[dict]Get DRS object metadata for a FASTA file.
Parameters:
digest(str): The sequence collection digest (which is also the DRS object ID)
Returns:
- dict: DRS object with id, self_uri, size, checksums, access_methods, etc.
service_info
Section titled “service_info”def service_info() -> Optional[dict]Get DRS service info.
Returns:
- dict: The service information.
write_fai
Section titled “write_fai”def write_fai(digest: str, dest_path: str, seqcol_client: 'SequenceCollectionClient' = None) -> strWrite a .fai index file for a FASTA.
Parameters:
digest(str): The sequence collection digestdest_path(str): Path to write the .fai fileseqcol_client(SequenceCollectionClient): SequenceCollectionClient to use (default:None)
Returns:
- str: Path to the written file
PangenomeClient
Section titled “PangenomeClient”class PangenomeClientBases: RefgetClient
Agent Classes
Section titled “Agent Classes”Agents provide higher-level abstractions for working with refget data in a PostgreSQL database.
RefgetDBAgent
Section titled “RefgetDBAgent”class RefgetDBAgent(engine: Optional[SqlalchemyDatabaseEngine] = None, postgres_str: Optional[str] = None, schema=SEQCOL_SCHEMA_PATH, inherent_attrs: List[str] = DEFAULT_INHERENT_ATTRS, fasta_drs_url_prefix: Optional[str] = None)Primary aggregator agent, interface to all other agents
Parameterized it via these environment variables:
- POSTGRES_HOST
- POSTGRES_DB
- POSTGRES_USER
- POSTGRES_PASSWORD
Attributes
Section titled “Attributes”Properties
Section titled “Properties”attribute
Section titled “attribute”@propertydef attribute -> AttributeAgentfasta_drs
Section titled “fasta_drs”@propertydef fasta_drs -> FastaDrsAgentpangenome
Section titled “pangenome”@propertydef pangenome -> PangenomeAgent@propertydef seq -> SequenceAgentseqcol
Section titled “seqcol”@propertydef seqcol -> SequenceCollectionAgentMethods
Section titled “Methods”aliases_for
Section titled “aliases_for”def aliases_for(kind: str, digest: str) -> listcalc_similarities
Section titled “calc_similarities”def calc_similarities(digestA: str, digestB: str) -> dictCalculates the Jaccard similarity between two sequence collections.
This method retrieves two sequence collections using their digests and then computes jaccard similarities for all attributes.
Parameters:
digestA(str): The digest (identifier) for the first sequence collection.digestB(str): The digest (identifier) for the second sequence collection.
Returns:
- dict: The Jaccard similarity score between the two sequence collections for all present and shared attributes.
calc_similarities_seqcol_dicts
Section titled “calc_similarities_seqcol_dicts”def calc_similarities_seqcol_dicts(seqcolA: dict, seqcolB: dict) -> dictCalculates the Jaccard similarity between two sequence collections.
This method retrieves one sequence collections using a digests and then computes jaccard similarities versus another input sequence collection dictionary.
Parameters:
seqcolA(dict): the first sequence collection in dict format.seqcolB(dict): the second sequence collection in dict format.
Returns:
- dict: The Jaccard similarity score between the two sequence collections for all present and shared attributes.
capabilities
Section titled “capabilities”def capabilities() -> dictcollection_count
Section titled “collection_count”def collection_count() -> intcompare_1_digest
Section titled “compare_1_digest”def compare_1_digest(digestA: str, seqcolB: dict) -> dictcompare_digest_with_level2
Section titled “compare_digest_with_level2”def compare_digest_with_level2(digest: str, level2_b: dict) -> dictcompare_digests
Section titled “compare_digests”def compare_digests(digestA: str, digestB: str) -> dictget_attribute
Section titled “get_attribute”def get_attribute(attribute_name: str, attribute_digest: str) -> listget_collection
Section titled “get_collection”def get_collection(digest: str, level: int = 2) -> dictget_collection_attribute
Section titled “get_collection_attribute”def get_collection_attribute(digest: str, attribute: str) -> listget_collection_itemwise
Section titled “get_collection_itemwise”def get_collection_itemwise(digest: str, limit: int | None = None) -> list[dict]get_fhr
Section titled “get_fhr”def get_fhr(digest: str)list_alias_namespaces
Section titled “list_alias_namespaces”def list_alias_namespaces(kind: str) -> listlist_aliases
Section titled “list_aliases”def list_aliases(kind: str, namespace: str) -> listlist_attributes
Section titled “list_attributes”def list_attributes(attribute: str, page: int = 0, page_size: int = 100) -> dictlist_collections
Section titled “list_collections”def list_collections(page: int = 0, page_size: int = 100, filters: dict | None = None) -> dictlist_fhr
Section titled “list_fhr”def list_fhr() -> listresolve_alias
Section titled “resolve_alias”def resolve_alias(kind: str, namespace: str, alias: str)retrieve_level2_digest
Section titled “retrieve_level2_digest”def retrieve_level2_digest(seqcoldigest: str) -> dicttruncate
Section titled “truncate”def truncate() -> intDelete all records from the database
SequenceCollectionAgent
Section titled “SequenceCollectionAgent”class SequenceCollectionAgent(engine: SqlalchemyDatabaseEngine, inherent_attrs: Optional[List[str]] = None, parent: Optional['RefgetDBAgent'] = None)Agent for interacting with database of sequence collection
Attributes
Section titled “Attributes”Methods
Section titled “Methods”def add(seqcol: SequenceCollection, update: bool = False) -> SequenceCollectionAdd a sequence collection to the database or update it if it exists
Parameters:
seqcol(SequenceCollection): The sequence collection to addupdate(bool): If True, update an existing collection if it exists (default:False)
Returns:
- SequenceCollection: The added or updated sequence collection
add_from_dict
Section titled “add_from_dict”def add_from_dict(seqcol_dict: dict, update: bool = False) -> SequenceCollectionAdd a sequence collection from a seqcol dictionary
Parameters:
seqcol_dict(dict): The sequence collection in dictionary formupdate(bool): If True, update an existing collection if it exists (default:False)
Returns:
- SequenceCollection: The added or updated sequence collection
add_from_fasta_file
Section titled “add_from_fasta_file”def add_from_fasta_file(fasta_file_path: str, update: bool = False, create_fasta_drs: bool = True, human_readable_name: str = None) -> SequenceCollectionGiven a path to a fasta file, load the sequences into the refget database.
Parameters:
fasta_file_path(str): Path to the fasta fileupdate(bool): If True, update an existing collection if it exists (default:False)create_fasta_drs(bool): If True, create a FastaDrsObject for the FASTA file (default:True)human_readable_name(str): Optional human-readable name for the collection (default:None)
Returns:
- SequenceCollection: The added or updated sequence collection
add_from_fasta_file_with_name
Section titled “add_from_fasta_file_with_name”def add_from_fasta_file_with_name(fasta_file_path: str, human_readable_name: str, update: bool = False, create_fasta_drs: bool = True) -> SequenceCollectionGiven a path to a fasta file, and a human-readable name, load the sequences into the refget database.
Deprecated: Use add_from_fasta_file(fasta_file_path, human_readable_name=name) instead.
add_from_fasta_pep
Section titled “add_from_fasta_pep”def add_from_fasta_pep(pep: 'peppy.Project', fa_root: str, update: bool = False, create_fasta_drs: bool = True) -> dictGiven a PEP project and a root directory containing the fasta files, load the fasta files into the refget database.
Parameters:
pep(peppy.Project): PEP project object containing sample metadatafa_root(str): Root directory containing the fasta filesupdate(bool): If True, update existing sequence collections (default:False)create_fasta_drs(bool): If True, create FastaDrsObjects for the FASTA files (default:True)
Returns:
- dict: A dictionary of the digests of the added sequence collections
def get(digest: str, return_format: str = 'level2', attribute: Optional[str] = None, itemwise_limit: Optional[int] = None) -> SequenceCollection | dict | listGet a sequence collection by digest
Parameters:
digest(str): The digest of the sequence collectionreturn_format(str): The format in which to return the sequence collection (default:'level2')attribute(str): Name of an attribute to return, if you just want an attribute (default:None)itemwise_limit(int): Limit the number of items returned in itemwise format (default:None)
Returns:
- SequenceCollection: The sequence collection (in requested format)
get_many_level2_offset
Section titled “get_many_level2_offset”def get_many_level2_offset(limit: int = 50, offset: int = 0, target_digests: Optional[List[str]] = None) -> ResultsSequenceCollectionsdef list(page_size: int = 100, cursor: Optional[str] = None) -> dictlist_by_offset
Section titled “list_by_offset”def list_by_offset(limit: int = 50, offset: int = 0) -> dictsearch_by_attributes
Section titled “search_by_attributes”def search_by_attributes(filters: dict, offset: int = 0, limit: int = 50) -> dictSearch sequence collections by multiple attribute filters (AND logic).
Parameters:
filters(dict): Dict of {attribute_name: digest} pairsoffset(int): Pagination offset (default:0)limit(int): Max results to return (default:50)
Returns:
- dict: Dict with pagination info and results
SequenceAgent
Section titled “SequenceAgent”class SequenceAgent(engine: SqlalchemyDatabaseEngine)Agent for interacting with database of sequences
Attributes
Section titled “Attributes”Methods
Section titled “Methods”def add(sequence: Sequence) -> Sequencedef get(digest: str, start: int | None = None, end: int | None = None) -> strdef list(offset: int = 0, limit: int = 50) -> dictPangenomeAgent
Section titled “PangenomeAgent”class PangenomeAgent(parent: 'RefgetDBAgent')Agent for interacting with database of pangenomes
Attributes
Section titled “Attributes”Methods
Section titled “Methods”def add(pangenome: Pangenome) -> Pangenomeadd_from_fasta_pep
Section titled “add_from_fasta_pep”def add_from_fasta_pep(pep: 'peppy.Project', fa_root: str, update: bool = False) -> Pangenomedef get(digest: str, return_format: str = 'level2') -> Pangenome | dictlist_by_offset
Section titled “list_by_offset”def list_by_offset(limit: int = 50, offset: int = 0) -> dictAttributeAgent
Section titled “AttributeAgent”class AttributeAgent(engine: SqlalchemyDatabaseEngine)Attributes
Section titled “Attributes”Methods
Section titled “Methods”def get(attribute_type: str, digest: str) -> listdef list(attribute_type: str, offset: int = 0, limit: int = 50) -> dictsearch
Section titled “search”def search(attribute_type: str, digest: str, offset: int = 0, limit: int = 50) -> dictFastaDrsAgent
Section titled “FastaDrsAgent”class FastaDrsAgent(engine: SqlalchemyDatabaseEngine, url_prefix: Optional[str] = None)Agent for interacting with database of FASTA DRS objects
Attributes
Section titled “Attributes”Methods
Section titled “Methods”def add(fasta_drs: FastaDrsObject) -> FastaDrsObjectAdd a FastaDrsObject to the database
add_access_method
Section titled “add_access_method”def add_access_method(digest: str, access_method: AccessMethod) -> FastaDrsObjectAdd an access method to an existing FastaDrsObject.
Parameters:
digest(str): The digest (object_id) of the DRS objectaccess_method(AccessMethod): The AccessMethod to add
Returns:
- FastaDrsObject: The updated FastaDrsObject
def get(digest: str) -> FastaDrsObjectGet a FastaDrsObject by its digest (object_id)
list_by_offset
Section titled “list_by_offset”def list_by_offset(limit: int = 50, offset: int = 0) -> dictList FastaDrsObjects with pagination
RefgetStore (gtars)
Section titled “RefgetStore (gtars)”RefgetStore provides high-performance local sequence storage implemented in Rust. It supports:
- In-memory and on-disk storage with optional compression
- Remote store access with local caching
- Sequence retrieval by digest or by collection + name
- BED file region extraction for batch operations
- FASTA export for individual sequences or regions
See the RefgetStore tutorial for usage examples.
RefgetStore, digest_fasta, compute_fai, digest_sequence, SequenceCollection, StorageMode, and sha512t24u_digest are implemented in Rust in gtars and re-exported by refget. Their API reference is in the gtars docs: gtars refget API reference.
Digest Functions
Section titled “Digest Functions”Low-level functions for computing GA4GH digests:
canonical_str
Section titled “canonical_str”def canonical_str(item: dict) -> bytesConvert a dict into a canonical string representation