Summary

Samenvatting Deeltoets 2 genomica

Rating

Sold

Pages

Uploaded on

28-09-2022

Written in

2021/2022

Alle hoorcolleges die bij het tweede deel van de cursus genomica gegeven worden zijn samengevat in dit document. Inclusief plaatjes, voorbeeld vragen van het college en formules.

Institution

Course

Content preview

Genomica – bioinformatica DT-2

Hoorcollege 1 - BLAST

Recap: -omics (omics: het sequencen van alles van iets) (meta: kijken naar alle organismen ipv 1)
- genomics: sequence all of the DNA of one organism
- transcriptomics: sequence all of the mRNA in an organism/tissue/cell
- proteomics: sequence all of the proteins in an organism/tissue/cell
- metagenomics: sequence the DNA of all organisms in a sample
- metatranscriptomics: sequence the mRNA of all organisms in a sample
- metaproteomics: sequence the proteins of all organisms in a sample

Hoe werkt metagenomics:
- de pakt een sample (bijv. koraal, zeewater, stuk darm, etc)
- filteren zodat je dingen kwijtraakt waar je niet naar wilt kijken
- dan hou je alleen de micro-organismen over

The biology behind the omics revolution
- omics solves a major problem in the science: biases
- people are mostly interested in: their diseases, their
food, themselves
- this causes biases in our general understanding of
biology, and biases in our databases. For example: most
studied bacteria are associated with humans

Data and the bioinformatician
- bioinformaticians use data in two different ways:
- 1: question first / top down: given a biological question, a good bioinformatician will immediately
think about which datasets could be used to answer it
- 2: data fist/ bottom up: given a dataset, a good bioinformatician will immediately think about which
biological hypothesis it could help to test

Bioinformatics
- the study of informatic process in biotic systems

Bioinformatic data analysis
- using computational methods to analyze biological data

What to do with tons of different sequences?
- searching a database: we want to find a query sequence in the database
- but why? → if two sequences are similar we assume that they are related or have a common
ancestor
- show database of 300k genomes and illustrate how we want to find the best hit of a given query
- we have to break down the search because of possible mutations

,k-mer searches
- sequences can be divided into shorter subsequences or k-mers (k-mers consist of k nucleotides or
amino acids)
- we can make an index of all k-mers that occur in the database sequences
- if we split a query into k-mers of the same length, we can rapidly identify all the database
sequences containing them
- but: we limit ourselves to exact matches

natural sequence divergence
- if we align metagenomics sequencing reads to a reference genome, we can distinguish multiple
distinct SAR86 strains
- the sequences at the top (~97% identity)
belong to a strain that is closely related to the
reference genome
- the sequences below (~60 – 80% identity) are
more distantly related strains

pairwise sequence alignments
- given two sequences: seqX = X1X2…XM and seqY = Y1Y2….YN
an alignment is an assignment of gaps to positions 0, …, M in x, and to positions 0,…,N in seqY, so as
to line up each letter in one sequence with either a letter of a gap in the other sequence
- je zet de sequences boven elkaar zodat er zo veel mogelijk overeenkomsten zijn

searching a database
- could we make sequence alignments between the query and every database sequence? →
theoratically, yes but it takes a long time

best of both worlds
- using an k-mer search (= index search) will be very fast… but limits you to the exact matches
- making all possible pairwise alignments will let you find distantly related sequences as well …. But it
would take a very long time
- the solution is to combine the best of both worlds: quickly find potential hit using exact k-mers
stored in an index and make pairwise alignment, but only for potential hits

basic local alignment search tool (BLAST)
- BLAST finds similar sequences at reasonable speed (10-50x faster than previous algorithms)
- terminology: query – sequence we search the database with. Hit or subject: similar sequence found
in the database
- BLAST is the most used bioinformatics program → more than 100.000 queries per day on the NCBI
BLAST server
- even faster algorithms are now available

the BLAST search algorithm
- 1: identifies all words (length W) in the query (default lengths: W = 3 for protein, W = 11 for DNA,
based on substitution scores)
- 2: quickly finds similar words in the database (similar words are defined by using the substitution

, matrix, the index quickly locates all potential hits seqs
- 3: extends seeds in both directions to find HSPs between query and hit (HSP: region that can be

aligned with a score above a certain threshold

Report Copyright Violation

Written for

Institution: Universiteit Utrecht (UU)
Study: Biologie
Course: Genomica

All documents for this subject (28)

Document information

Uploaded on: September 28, 2022
Number of pages: 28
Written in: 2021/2022
Type: SUMMARY

Subjects

genomica
lena will
utrecht university
biologie
bioinformatica
universiteit utrecht
uu

$7.16

Get access to the full document:

Written by students who passed

Immediately available after payment

Read online or as PDF

Get to know the seller

charlottebruring

Get to know the seller

charlottebruring Universiteit Utrecht

View profile

Sold

Member since

3 year

Number of followers

Documents

Last sold

0.0

0 reviews

Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Frequently asked questions

What do I get when I buy this document?

You get a PDF, available immediately after your purchase. The purchased document is accessible anytime, anywhere and indefinitely through your profile.

Satisfaction guarantee: how does it work?

Our satisfaction guarantee ensures that you always find a study document that suits you well. You fill out a form, and our customer service team takes care of the rest.

Who am I buying these notes from?

Stuvia is a marketplace, so you are not buying this document from us, but from seller charlottebruring. Stuvia facilitates payment to the seller.

Will I be stuck with a subscription?

No, you only buy these notes for $7.16. You're not tied to anything after your purchase.

Can Stuvia be trusted?

4.6 stars on Google & Trustpilot (+1000 reviews) 57108 documents were sold in the last 30 days Founded in 2010, the go-to place to buy study notes for 16 years now

Samenvatting Deeltoets 2 genomica

Content preview

Written for

Document information

Subjects

Get to know the seller

Recently viewed by you

Why students choose Stuvia

Created by fellow students, verified by reviews

Didn't get what you expected? Choose another document

Pay as you like, start learning right away

Working on your references?

Frequently asked questions

What do I get when I buy this document?

Satisfaction guarantee: how does it work?

Who am I buying these notes from?

Will I be stuck with a subscription?

Can Stuvia be trusted?