Loading UTA

Loading is largely driven by the Makefile in loading/Makefile. You don't
have to do it this way, but the details of the loading appear there.


Loading occurs in three distinct stages:
  1. extraction and translations from data sources into intermediate files
  2. loading from intermediates into UTA.
  3. generating new exon alignments

The extraction scripts, which are specific for each source, are in
uta/sbin/. Each script reads from one source and writes one or more
intermediates in formats that are specified in uta/formats/.  The
intermediate file formats are indpendent of source.  See
uta/doc/misc-figures.pdf.


How Reece does it
-----------------

Updated 2014-08-29
Tested with Ubuntu 14.04, x86_64, PostgreSQL 9.3.4

*) get sequence fasta files

Loading requires a lot of fasta files to be available.  Non-Invitae
users should populate a directory with fasta files with all sequences
you intend to load. Typically, this includes chromosome sequence fasta
files (with NC annotations) and RefSeq fasta files (with NM at least).
Possible sources are ftp://ftp.ncbi.nlm.nih.gov/genomes/H_sapiens/
(only fa.gz files are required) and
ftp://ftp.ncbi.nlm.nih.gov/refseq/H_sapiens/mRNA_Prot/ (only .faa.gz
and .fna.gz).  Then, update uta/etc/global.conf with the path to the
downloads.

Invitae users may wish to rsync /locus/data/core-rest/2014-01/ to
their machine (6.7GB).  That directory includes Ensembl, BIC,
RefSeqGene, and older NCBI RefSeqs.


*) install postgresql

Source or distribution.


*) Create the uta_dev database

$ export PGDATA=/local/home/reece/var/pg/9.3-test
$ export PGPORT=5433
$ export PGDATABASE=uta_dev

# PGDATA and PGPORT are listed on the command lines below to be
# explicit but are not strictly required because postgresql tools
# recognize those variables.

$ initdb -D $PGDATA
$ perl -i.bak -pe "s/#port = 5432/port = $PGPORT/" $PGDATA/postgresql.conf
$ pg_ctl -D $PGDATA -l logfile start
(trust authentication enabled and required for these instructions)


*) prepare a virtualenv

$ mkvirtualenv uta


*) prepare the uta working directory

$ hg clone ssh://hg@bitbucket.org/uta/uta uta_dev
$ cd uta_dev
$ make develop
TODO: error the first time (re: distribute package). Second time was clean. Figure this out.
$ make develop


*) install non-public tools manually

pip install hg+ssh://hg@bitbucket.org/locusdevelopment/locus-core
pip install hg+ssh://hg@bitbucket.org/locusdevelopment/locus-lib-bio

NOTE: Reece is currently (2014-09-03) working on extracting the
essential elements of these repos into a public package.  To be
discussd with Geoff prior to release.


*) prepare a database

$ createuser uta_admin
$ createuser uta_public
$ createdb -O uta_admin uta_dev


*) create and load a test database

The general command to load data is
$ make db-build DATA=<dataset> CONF=<confname>

For testing, type

$ make db-build DATA=test CONF=test

DATA=test refers to the directory uta/loading/test for source data
CONF=test refers to the file ../etc/test.conf for connection parameters.

uta/loading/test contains excerpts of a full load. Those files are
part of the repo to enable easy testing of loading.

If that fails, 

$ psql -p $PGPORT -d <db> -c 'drop schema uta1 cascade'
$ make cleanest

and retry.


## Everything past this is for a full load


*) Extraction and translation

$ make main-data

This will take a long time. 36 hours maybe.


*) uncompress resulting fasta files into the fasta directory (see
main.conf)


*) Optional: make test data
These data are used for testing and therefore committed with the repo.
You probably don't need to rebuild them.

make test-data


*) create and load a database

$ make db-build DATA=main CONF=test


*) push to RDS

$ make push


