Skip to main content

Genes and Genomics

Overview

  • Credit value: 15 credits at Level 6
  • Convenor: to be confirmed
  • Assessment: in-class problem sets (20%), a computer lab practical (30%) and an end-of-module test (50%)

Module description

Biosciences now rely heavily on biological databases, the analysis of sequences and structures, and the interpretation of large datasets. In this module we introduce you to a range of core bioinformatics resources popular across biosciences. We will also equip you with the practical bioinformatics skills required to make use of these resources and to analyse and interpret biological data.

Indicative syllabus

  • Genes and genomes: What is a gene?, What is homology?, How are genomes organised in prokaryotes and eukaryotes?, What databases are available to access genomic data?
  • Analysis of biological sequences (I): assessing pairwise sequence similarity; using heuristic approaches (BLAST) to search a database for homologues using a query sequence
  • Analysis of biological sequences (II): multiple sequence alignments and models of protein families; additional uses of multiple sequence alignments (regulatory motifs, sequence logos)
  • Accessing and analysing protein structures: introduction to the PDB database, understanding the information held
  • Measuring gene expression across a whole genome; introduction to NGS-based transcriptomic data and databases (GEO)
  • Protein and gene networks; introduction to network databases (STRING)
  • From next-generation sequencing reads to genomes: reconstructing and annotating small genomes from genomic reads deposited in public databases using the BV-BRC server
  • Comparing genomes: variants, disease and personalised medicine: introduction to databases holding genetic variation in humans (ClinVar, COSMIC); exploring web servers that predict the effect of such variations; results of GWAS studies; ethical considerations

Learning objectives

By the end of this module, you will be able to:

  • identify and explain key features of genome organisation in prokaryotic and eukaryotic organisms
  • navigate confidently protein and DNA databases (Uniprot, NCBI, Ensembl) and retrieve sequences of individual proteins or genes; obtain and analyse search results from other popular databases used widely in bioinformatics (PDB for 3D structures, dbSNP/OMIM for mutations, GEO for gene expression data)
  • carry out sequence-based analyses using bioinformatics servers (e.g. EMBOSS tools or the NCBI servers): pairwise and multiple alignments of proteins/DNA, database searching for sequences similar to a query (BLAST/Hmmer)
  • understand what information is held in a PDB record of a 3D structure and critically assess the quality of a structure, given the metrics available in the record.
  • describe in simple terms next-generation sequencing for genomic and transcriptomic data; be familiar with the format of such data and how to access it in publicly available repositories (ENA, GEO, NCBI); use competently online servers to carry out simple differential expression comparisons between transcriptomic datasets in GEO
  • search for information protein/gene network databases (STRING) and critically evaluate the results returned by such searches
  • use the BV-BRC server to reconstruct a small bacterial or viral genome, using raw NGS reads downloaded from a public database
  • show awareness of limitations of using such servers and interpret the results of reconstructing and annotating the genome automatically
  • access databases holding variants (e.g. ClinVar) and retrieve information linking these variants to phenotypes or diseases
  • understand ethical issues associated with sequencing and storing human genomes and carrying out analyses of the variations between individuals and groups of individuals
  • integrate information from multiple sources to study a particular gene or protein: sequence, structure, interactions with small molecules and other proteins/genes, information on known variations and prediction of their effect, etc.