An ORFeome-based analysis of human transcription factor genes and the construction of a microarray to interrogate their expression
- PMID: 15489324
- PMCID: PMC528918
- DOI: 10.1101/gr.2584104
An ORFeome-based analysis of human transcription factor genes and the construction of a microarray to interrogate their expression
Abstract
Transcription factors (TFs) are essential regulators of gene expression, and mutated TF genes have been shown to cause numerous human genetic diseases. Yet to date, no single, comprehensive database of human TFs exists. In this work, we describe the collection of an essentially complete set of TF genes from one depiction of the human ORFeome, and the design of a microarray to interrogate their expression. Taking 1468 known TFs from TRANSFAC, InterPro, and FlyBase, we used this seed set to search the ScriptSure human transcriptome database for additional genes. ScriptSure's genome-anchored transcript clusters allowed us to work with a nonredundant high-quality representation of the human transcriptome. We used a high-stringency similarity search by using BLASTN, and a protein motif search of the human ORFeome by using hidden Markov models of DNA-binding domains known to occur exclusively or primarily in TFs. Four hundred ninety-four additional TF genes were identified in the overlap between the two searches, bringing our estimate of the total number of human TFs to 1962. Zinc finger genes are by far the most abundant family (762 members), followed by homeobox (199 members) and basic helix-loop-helix genes (117 members). We designed a microarray of 50-mer oligonucleotide probes targeted to a unique region of the coding sequence of each gene. We have successfully used this microarray to interrogate TF gene expression in species as diverse as chickens and mice, as well as in humans.
Figures
References
-
- Apweiler, R., Attwood, T.K., Bairoch, A., Bateman, A., Birney, E., Biswas, M., Bucher, P., Cerutti, L., Corpet, L., Croning, M.D.R., et al. 2001. The InterPro database, an integrated documentation resource for protein families, domains and functional sites. Nucleic Acids Res. 29: 37-40. - PMC - PubMed
-
- Boutanaev, A., Kalmykova, A.I., Shevelyov, Y.Y., and Nurminsky, D. 2002. Large clusters of co-expressed genes in the Drosophila genome. Nature 420: 666-669. - PubMed
-
- Boyadiev, S.A. and Jabs, E.W. 2000. Developmental biology: Frontiers for clinical genetics. Clin. Genet. 57: 253-266. - PubMed
-
- Brivanlou, A.H. and Darnell Jr., J.E. 2002. Signal transduction and the control of gene expression. Science 295: 813-818. - PubMed
WEB SITE REFERENCES
-
- http://www.ensembl.org; Ensembl genome browser.
-
- http://flybase.bio.indiana.edu/; FlyBase, a database of the Drosophila genome.
-
- http://www.geneontology.org/; Gene Ontology Consortium.
-
- http://www.ebi.ac.uk/interpro/; InterPro.
-
- http://www.ncbi.nlm.nih.gov/LocusLink/; LocusLink.
Publication types
MeSH terms
Substances
LinkOut - more resources
Full Text Sources
Other Literature Sources
Research Materials
Miscellaneous