0
settings
الوضع الليلي
moon
انماط الصفحة الرئيسية arrow
EN
1
المرجع الالكتروني للمعلوماتية

النبات

مواضيع عامة في علم النبات

الجذور - السيقان - الأوراق

النباتات الوعائية واللاوعائية

البذور (مغطاة البذور - عاريات البذور)

الطحالب

النباتات الطبية

الحيوان

مواضيع عامة في علم الحيوان

علم التشريح

التنوع الإحيائي

البايلوجيا الخلوية

الأحياء المجهرية

البكتيريا

الفطريات

الطفيليات

الفايروسات

علم الأمراض

الاورام

الامراض الوراثية

الامراض المناعية

الامراض المدارية

اضطرابات الدورة الدموية

مواضيع عامة في علم الامراض

الحشرات

التقانة الإحيائية

مواضيع عامة في التقانة الإحيائية

التقنية الحيوية المكروبية

التقنية الحيوية والميكروبات

الفعاليات الحيوية

وراثة الاحياء المجهرية

تصنيف الاحياء المجهرية

الاحياء المجهرية في الطبيعة

أيض الاجهاد

التقنية الحيوية والبيئة

التقنية الحيوية والطب

التقنية الحيوية والزراعة

التقنية الحيوية والصناعة

التقنية الحيوية والطاقة

البحار والطحالب الصغيرة

عزل البروتين

هندسة الجينات

التقنية الحياتية النانوية

مفاهيم التقنية الحيوية النانوية

التراكيب النانوية والمجاهر المستخدمة في رؤيتها

تصنيع وتخليق المواد النانوية

تطبيقات التقنية النانوية والحيوية النانوية

الرقائق والمتحسسات الحيوية

المصفوفات المجهرية وحاسوب الدنا

اللقاحات

البيئة والتلوث

علم الأجنة

اعضاء التكاثر وتشكل الاعراس

الاخصاب

التشطر

العصيبة وتشكل الجسيدات

تشكل اللواحق الجنينية

تكون المعيدة وظهور الطبقات الجنينية

مقدمة لعلم الاجنة

الأحياء الجزيئي

مواضيع عامة في الاحياء الجزيئي

علم وظائف الأعضاء

الغدد

مواضيع عامة في الغدد

الغدد الصم و هرموناتها

الجسم تحت السريري

الغدة النخامية

الغدة الكظرية

الغدة التناسلية

الغدة الدرقية والجار الدرقية

الغدة البنكرياسية

الغدة الصنوبرية

مواضيع عامة في علم وظائف الاعضاء

الخلية الحيوانية

الجهاز العصبي

أعضاء الحس

الجهاز العضلي

السوائل الجسمية

الجهاز الدوري والليمف

الجهاز التنفسي

الجهاز الهضمي

الجهاز البولي

المضادات الميكروبية

مواضيع عامة في المضادات الميكروبية

مضادات البكتيريا

مضادات الفطريات

مضادات الطفيليات

مضادات الفايروسات

علم الخلية

الوراثة

الأحياء العامة

المناعة

التحليلات المرضية

الكيمياء الحيوية

مواضيع متنوعة أخرى

الانزيمات

قم بتسجيل الدخول اولاً لكي يتسنى لك الاعجاب والتعليق.

Carrying out sequence assembly in complex genomes

المؤلف:  Strachan, T., & Read, A.

المصدر:  Human molecular genetics

الجزء والصفحة:  5th E, P218-219

2026-09-28

20

+

-

20

Early approaches to sequencing a complex genome involved hierarchical shotgun sequencing (see Figure 1B). The aim then was to identify the minimum number of DNA clones that need to be sequenced to provide an acceptable “sequence depth” (ide ally, a region of interest on the DNA molecule should be represented by a large number of sequence reads to maximize the chances of obtaining an unambiguous consensus sequence). The ideal template for genome assembly would be a continuous clone con tig for each chromosomal DNA molecule. That is, for each chromosome there would be a continuous tiling path of clones with overlapping DNA inserts. (Visualize the contig shown in Figure 2 extending across the whole chromosome.) For complex genomes, however, there are problems achieving that aim.

Fig1. Two strategies for sequencing a genome. (A) Whole-genome shotgun sequencing involves indiscriminate fragmentation of the genome into small pieces of DNA that are readily sequenced. It quickly generates large amounts of sequence data but anchoring sequences to specific locations in the genome may be problematic. Large amounts of repetitive DNA in complex genomes make it difficult to unambiguously locate sequences to specific subchromosomal regions. (B) For first-time sequencing of complex genomes, it is more efficient to assemble contigs of large insert clones for each chromosome and then to fragment individual clones into pieces that are sequenced to reconstruct the sequence of the parent clone. (Adapted from Waterston RH et al. [2002] Proc Natl Acad Sci USA 99:3712–3716; PMID 11880605. With permission from National Academy of Sciences. Copyright [2002] National Academy of Sciences, USA.)

Fig2. Assembling clones in a clone contig. Because partial digestion of genomic DNA is employed to construct a genomic DNA library, some DNA clones from a genome region will have inserts that partially overlap with each other. Here, the chromosomal DNA sequence from positions A to B is represented by a linear series of overlapping DNA inserts, a tiling path. The collection of clones whose inserts produce a tiling path is known as a clone contig.

The long arrays of tandem, highly-repetitive DNA sequences associated with constitutive heterochromatin provide a major obstacle for genome assembly: very high levels of sequence identity between the repeats in a long array mean that identifying overlapping sequences is difficult. Some gene clusters provide obstacles to genome assembly, too, such as the long arrays of tandem repeats specifying the 18S, 5.8S, and 28S ribosomal RNAs.

Outside difficulties with long arrays of tandem repeats, different problems can affect genome assembly. Some genome regions may not be represented by sequence reads (certain sequences may not be propagated well in bacteria, for example, and so be poorly represented in libraries of DNA clones). Challenges can also be posed by other, closely related repeats. And structural variation between haplotypes can pose problems for establishing contigs. As a result of these difficulties, the sequences of even the euchromatic part of complex genomes, such as the human genome, are unfinished. A chromosome is typically represented by scaffolds that each contain multiple contigs, with individual contigs separated by gaps whose approximate lengths are known (Figure 3).

Fig3. Scaffolds and contigs. Contigs are made up of overlapping DNA sequences without gaps. Scaffolds are made up of two or more contigs with gaps, where the gaps are of approximately known length and each contig has a known orientation within the scaffold. (A) An example of a scaffold with three clone contigs developed in a hierarchical shotgun sequencing strategy. The blue bars represent overlapping sequenced inserts of large-insert clones (whose sequences had been obtained after shotgun sequencing). (B) An example of contigs and scaffolds established using massively parallel sequencing of paired ends (pale green and orange boxes) and mate-pair sequences (dark green and red boxes). The approximate size of the large gap between contigs Y and Z may be known because of the linked mate-pair sequences (connected by dashed blue lines) that were originally derived from sequences that can be up to tens of kilobases apart in the genome (but have been artificially brought together onto short fragments for sequencing.)

Assembly statistics

Different assembly statistics are used. N50, the most commonly used statistic, describes an average length of a set of sequences (contigs or scaffolds), but it is not the mean or median length. It is defined as the largest length L such that 50% of all nucleotides are contained in contigs of size at least L. To calculate a contig N50, for example, every contig is first ordered by length from longest to shortest. Then, starting from the longest contig, the lengths of each contig are sequentially summed until this running sum equals one half of the total length of all contigs in the assembly. At that point, the length of the shortest contig in the list is the contig N50. L50, another commonly used statistic, signifies the number of sequences whose sum length exceeds 50% of the total size of the sequence assembly. (Note: because of the confusing initial letters, some people have preferred to invert the meanings of N50 and L50, so that N50 becomes a number and L50 represents a length.) Table 1 provides an example of statistics for scaffolds and contigs for the human genome reference sequence.

Table1. SOME ASSEMBLY STATISTICS FOR THE HUMAN GENOME REFERENCE SEQUENCE GRCh38.p12

اشترك بقناتنا على التلجرام ليصلك كل ما هو جديد