- 论坛徽章:
- 1
|
有一些氨基酸序列,文件名为a.fas,格式如下:>Gne|c67372_g2_i2_m.86976GIGGGLSGTVRQKGSIAPPPASGRIRSPLPPPPNDTVTIR---------STD-ARHTTDAFSDLSQLERNLPAPTASGWAAF
>Lja|Lygodium_c28780_g1_i1_m31025
SGKSSLLSTGHMAAPLAPPPGGGRLRTPLPPPPNDIKHHRVSNSTPLTRDTASGRP-GDPFADLSQLEANLPTSYHAGWAAF
>Lja|Lygodium_c30483_g1_i1_m35786
SGKSSLLSTGHMAAPLAPPPGGGRLRTPLPPPPNDIKHHRVSNSTPLTRDTASGRP-GDPFADLSQLEANLPTSYHAGWAAF
>Lja|Lygodium_c41319_g1_i1_m77659
GEQTNIASTGRMKAPLAPPPGGGRLRSPLPPPPSDNRQLR-SSSAANTYTEKTAKHASDPFADIAELETSLPSSFHGGWAAF
>Paq|Pteridium_c51474_g1_i3_m67055
SGGSTLASPGSMKAPIAPPPSGGRLCSPLPPPPNDTKYLRVTRTLPSTQ---------------------------------
>Paq|Pteridium_c62366_g1_i2_m115536
SGGSTLASPGSMKAPIAPPPSGGRLCSPLPPPPNDTKYLRVTRTLPSTQSSHSGRPTGDPFADLSQLEELLCARQHFEFASF
>Paq|Pteridium_c64930_g1_i1_m134191
GGTSTVASSGHIKAPLAPPPGGGRLRSPLPPPPNDNKHSRGTNAAAASQSRDSGRPTGDPFADLSQLEVSLPSSLHAGWAAF
>Pin|c31137_g1_i1_m49244
GLAGSVSSTGRQKASLAPPPGSGRIRSPLPPPPNDTVTAKIGSAMISSSSRDTSRHNADPLSDLSQLEMSLPSSTASGWAAF
其中|线前面的为物种名,基本都是3个字母。目的是想如果物种名相同,如果有多条序列,相差较大的话,所有序列都留下;如果排列好的序列相同或相差不多于2个氨基酸,则删除较短的一条。如上红色所示,将粉色的其中任意一条去除,红色的去除第一条较短的序列,得到如下结果(a.fas_out):
>Gne|c67372_g2_i2_m.86976GIGGGLSGTVRQKGSIAPPPASGRIRSPLPPPPNDTVTIR---------STD-ARHTTDAFSDLSQLERNLPAPTASGWAAF
>Lja|Lygodium_c28780_g1_i1_m31025
SGKSSLLSTGHMAAPLAPPPGGGRLRTPLPPPPNDIKHHRVSNSTPLTRDTASGRP-GDPFADLSQLEANLPTSYHAGWAAF
>Lja|Lygodium_c41319_g1_i1_m77659
GEQTNIASTGRMKAPLAPPPGGGRLRSPLPPPPSDNRQLR-SSSAANTYTEKTAKHASDPFADIAELETSLPSSFHGGWAAF
>Paq|Pteridium_c62366_g1_i2_m115536
SGGSTLASPGSMKAPIAPPPSGGRLCSPLPPPPNDTKYLRVTRTLPSTQSSHSGRPTGDPFADLSQLEELLCARQHFEFASF
>Paq|Pteridium_c64930_g1_i1_m134191
GGTSTVASSGHIKAPLAPPPGGGRLRSPLPPPPNDNKHSRGTNAAAASQSRDSGRPTGDPFADLSQLEVSLPSSLHAGWAAF
>Pin|c31137_g1_i1_m49244
GLAGSVSSTGRQKASLAPPPGSGRIRSPLPPPPNDTVTAKIGSAMISSSSRDTSRHNADPLSDLSQLEMSLPSSTASGWAAF
请大家帮帮忙!!!非常感谢!!!
|
|