Gene EcSMS35_4740 details

Gene Information       Plasmid Coverage information       Fosmid Coverage information       Sequence       

Gene Information

Locus tagEcSMS35_4740 
SymbolpepA 
ID6147238 
TypeCDS 
Is gene splicedNo 
Is pseudo geneNo 
Organism nameEscherichia coli SMS-3-5 
KingdomBacteria 
Replicon accessionNC_010498 
Strand
Start bp4839389 
End bp4840900 
Gene Length1512 bp 
Protein Length503 aa 
Translation table11 
GC content55% 
IMG OID641619555 
Productleucyl aminopeptidase 
Protein accessionYP_001746663 
Protein GI170681979 
COG category[E] Amino acid transport and metabolism 
COG ID[COG0260] Leucyl aminopeptidase 
TIGRFAM ID 


Plasmid Coverage information

Num covering plasmid clones
Plasmid unclonability p-value0.00000218455 
Plasmid hitchhikingYes 
Plasmid clonabilityhitchhiker 
 

Fosmid Coverage information

Num covering fosmid clones46 
Fosmid unclonability p-value0.712457 
Fosmid HitchhikerNo 
Fosmid clonabilitynormal 
 

Sequence

Gene sequence
ATGGAGTTTA GTGTAAAAAG CGGTAGCCCG GAGAAACAGC GGAGTGCCTG CATCGTCGTG 
GGCGTCTTCG AACCACGTCG CCTTTCTCCG ATTGCAGAAC AGCTCGATAA AATCAGCGAT
GGGTACATCA GCGCCCTGCT ACGTCGGGGC GAACTGGAAG GAAAACCGGG GCAGACACTG
TTGCTGCACC ATGTTCCGAA TGTACTTTCC GAGCGAATTC TCCTTATTGG TTGCGGCAAA
GAACGTGAGC TGGATGAACG TCAGTACAAG CAGGTTATTC AGAAAACCAT TAATACGCTG
AATGATACTG GCTCAATGGA AGCGGTCTGC TTTCTGACTG AACTGCACGT TAAAGGCCGT
AACAACTACT GGAAAGTGCG TCAGGCTGTC GAGACTGCAA AAGAGACGCT CTACAGTTTC
GATCAGCTGA AAACGAACAA GAGCGAACCG CGTCGTCCGC TGCGTAAAAT GGTGTTCAAC
GTGCCGACCC GCCGTGAACT GACCAGCGGT GAGCGCGCGA TCCAGCACGG TCTGGCGATT
GCCGCCGGGA TTAAAGCAGC AAAAGATCTC GGCAATATGC CGCCGAATAT CTGTAACGCC
GCTTACCTCG CTTCACAAGC GCGCCAGCTG GCTGACAGCT ACAGCAAGAA TGTCATCACC
CGCGTTATCG GCGAACAGCA GATGAAAGAG CTGGGGATGC ATTCTTATCT GGCGGTCGGT
CAGGGTTCGC AGAACGAATC GCTGATGTCG GTGATTGAGT ACAAAGGCAA CGCGTCGGAA
GATGCTCGCC CAATCGTGCT GGTGGGTAAA GGTTTAACCT TCGACTCCGG CGGTATCTCC
ATCAAGCCTT CAGAAGGCAT GGATGAGATG AAGTACGATA TGTGCGGCGC GGCGGCGGTT
TACGGCGTGA TGCGTATGGT CGCGGAGCTG CAACTGCCGA TTAACGTTAT CGGCGTGCTG
GCAGGCTGCG AAAACATGCC TGGCGGGCGT GCCTATCGTC CGGGCGATGT GTTAACCACC
ATGTCCGGTC AAACCGTTGA AGTGCTGAAC ACCGATGCCG AAGGCCGCCT GGTACTGTGC
GACGTGTTAA CTTACGTTGA ACGTTTTGAG CCGGAAGCGG TGATTGATGT GGCGACGCTG
ACCGGTGCCT GCGTGATCGC GCTGGGTCAT CATATTACTG GTCTGATGGC GAACCATAAT
CCGCTGGCCC ATGAACTGAT TGCCGCGTCT GAACAATCCG GTGACCGCGC ATGGCGCTTA
CCGCTGGGTG ACGAGTATCA GGAACAACTG GAGTCCAATT TTGCCGATAT GGCGAACATT
GGCGGTCGTC CTGGTGGGGC GATTACCGCA GGTTGCTTCC TGTCACGCTT TACCCGTAAG
TACAACTGGG CGCACCTGGA TATTGCAGGA ACCGCCTGGC GTTCTGGTAA AGCAAAAGGC
GCAACCGGTC GTCCGGTAGC GTTGCTGGCA CAGTTCCTGC TGAATCGCGC TGGGTTTAAC
GGCGAAGAGT AA
 
Protein sequence
MEFSVKSGSP EKQRSACIVV GVFEPRRLSP IAEQLDKISD GYISALLRRG ELEGKPGQTL 
LLHHVPNVLS ERILLIGCGK ERELDERQYK QVIQKTINTL NDTGSMEAVC FLTELHVKGR
NNYWKVRQAV ETAKETLYSF DQLKTNKSEP RRPLRKMVFN VPTRRELTSG ERAIQHGLAI
AAGIKAAKDL GNMPPNICNA AYLASQARQL ADSYSKNVIT RVIGEQQMKE LGMHSYLAVG
QGSQNESLMS VIEYKGNASE DARPIVLVGK GLTFDSGGIS IKPSEGMDEM KYDMCGAAAV
YGVMRMVAEL QLPINVIGVL AGCENMPGGR AYRPGDVLTT MSGQTVEVLN TDAEGRLVLC
DVLTYVERFE PEAVIDVATL TGACVIALGH HITGLMANHN PLAHELIAAS EQSGDRAWRL
PLGDEYQEQL ESNFADMANI GGRPGGAITA GCFLSRFTRK YNWAHLDIAG TAWRSGKAKG
ATGRPVALLA QFLLNRAGFN GEE