Structural Insights and Ligand Binding Site Analysis of Prolyl Endo-Protease (PEP): A Promising Insecticide Targeting Eurygaster integriceps (Sunn Pest)
Effat Nooriï¿¿ 1,2
Mojgan Bandehpour 2,3
Bahram Kazemi 2,3
Fatemeh Sefid 4
Sajad Najafi 2
Ali Ahmadizad Firouzjaei 2
Effat Noori 1,5✉ Email
1
A
Department of Medical Biotechnology, Faculty of Medicine Shahed University Tehran Iran
2
A
Department of Medical Biotechnology, School of Advanced Technologies in Medicine Shahid Beheshti University of Medical Sciences Tehran Iran
3
A
Cellular and Molecular Biology Research Center Shahid Beheshti University of Medical Sciences Tehran Iran
4 Department of Medical Genetics Shahid Sadoughi University of Medical Science Yazd Iran
5
A
School of Advanced Technologies in Medicine Shahid Beheshti University of Medical Sciences Tehran Iran
6
A
A
0098-021 51212600
Effat Noori∗ 1,2 , Mojgan Bandehpour2,3, Bahram Kazemi2,3, Fatemeh Sefid4, Sajad Najafi2, Ali Ahmadizad Firouzjaei2
1 Department of Medical Biotechnology, Faculty of Medicine, Shahed University, Tehran, Iran
2Department of Medical Biotechnology, School of Advanced Technologies in Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran
3Cellular and Molecular Biology Research Center, Shahid Beheshti University of Medical Sciences, Tehran, Iran
4 Department of Medical Genetics, Shahid Sadoughi University of Medical Science, Yazd, Iran
*Correspond author: Effat Noori, effat.noori@yahoo.com
Department of Medical Biotechnology, Faculty of Medicine, Shahed University, Tehran, Iran
School of Advanced Technologies in Medicine, Shahid Beheshti University of Medical Sciences, Tehran, Iran
Tel: 0098 − 021 51212600
Abstract
Prolyl Endo-Protease (PEP) is one of the crucial enzymes found in the salivary gland of Eurygaster integriceps. This study reveals PEP structure via in silico methods. BLAST was performed on the sequence to find the most appropriate template. The model of the 3D structure was made using the template and quality evaluation was performed for all models. The determination of ligand binding sites as well as the refinement of the 3D structure was performed. In the topology model presented here, this protein doesn’t have any transmembrane region. Five pockets on protein surfaces were obtained using the GHECOM server. COFACTOR software is used for finding ligand binding sites indicating the involvement of conserved residues, especially 173, 472, 475, 553, 554, 599, 644, 645, and 681 in the ligand binding site. Overall, this study provides detailed information, which can be helpful in designing highly efficient pesticides by inhibiting the Eurygaster integriceps proteases.
Keywords:
Eurygaster integriceps
Prolyl Endo-Protease
insecticide
in silico
Wheat
1. Introduction
The Eurygaster integriceps Puton )Sunn Pest( of the family Scutelleridae, is known as the most serious insect pest contaminating some essential food sources, including wheat and other cereal crops in many geographical areas, including Southern and Eastern Europe, Near East, and Pacific region (13). One of the current challenges in the world is the growing population and requirements for food security. It is estimated that by 2050 the world population will reach 9 billion. Wheat grain is recognized among the main foods for about 2 billion people all over the world (4). It is estimated that Sunn Pest infests over 15 million hectares of crops in the Middle Eastern countries and damages caused by these pests are reported 20–30% in barley and 50–90% in wheat. In the event that Sunn Pest infest is not controlled, the amount of damage can reach 100%. Both nymphs and adult forms of life identified for this pest are capable of causing a direct reduction in wheat yield by contaminating and feeding on various parts including grains, leaves, and stems (1, 2, 57).
The toxic impacts of Sunn Pests are partly due to the injection of their degrading enzymes found in salivary glands into the body of wheat plants, which causes considerable damage to the quality and yield of produced wheat (1, 8). These hydrolytic and proteolytic enzymes of the salivary gland remain in the grain. The texture of bread prepared from these affected seeds is weak, and sticky such that 2%–5% Sunn Pest-contaminated seeds lack the quality required for baking (1, 2, 57).
Serine, aspartic, cysteine, and metalloproteases are insect digestive proteases. Inhibition of these proteases can decrease the amount of several amino acids, which are essential for insect survival and growth development (9).
Sunn Pest's life cycle is composed of two major periods including active (feeding) and inactive (non-feeding) stages. This insect grows and develops during the active period in wheat fields and stores the amounts of energy required for survival in its inactive period. The devastating nature of damaging Sunn Pest in strategic plants, such as wheat makes pesticide application unavoidable (10, 11). Targeting the insect nutrition system is considered the best strategy for effective and specific control of this insect pest. (5).
Generally, the main management of the Sunn Pest is through using chemical material and biopesticides, natural enemies including parasitoids, digestive enzyme inhibitors, and insect-resistant genes in wheat plants (7).
One of the important enzymes of the salivary gland is prolyl endoprotease (PEP), which the first enzyme was identified by Darkoh et al. (5). This enzyme is a serine protease with the ability to digest intact wheat gluten (1, 8).
A
In our previous study, the PEP from the Sunn Pest was expressed in E. coli and showed the ability to cleave gluten specifically in vitro (12). Due to the significant diversity in Sunn pest digestive proteolytic enzymes, it is essential to study the vital enzymes of insects in detail to design a reasonable control strategy. Specific deactivation of digestive enzymes through using inhibitors in those insects results in malnourishment and eventually death from starvation (8).
Unfortunately, an increasing number of insecticide applications has exacerbated the emergence of pesticide resistance (7, 13). Resistant Sunn Pests to modern insecticides, mainly pyrethroids, have been reported in Russia (14).
In addition, studies have shown that many pests quickly develop resistance to chemical insecticides, including organic phosphor and carbamate (15). The computational methods have accelerated the analysis process and protein function prediction quickly and inexpensively (16). In silico annotation of proteinsʼ functional and structure with high precision has emerged as a crucial approach helpful for deep unveiling the molecular mechanism (17).
The first stage in pesticide design is in silico study to assess the competence of the candidate molecule. Taken together, in silico experiments have the potential to specify the target molecules for possible ligand binding sites, evaluate their capability to be considered as a drug or pesticide, and refine structures to improve binding characteristics (1820). Because of considerable challenges and the time-consuming identity of experimental analyses, (18), in the present study, we used bioinformatics instruments for function and 3D structure determination. This study aimed to characterize the PEP of Sunn Pest, which is most commonly found in their salivary glands secretions as well as damaged grains to provide more detailed information to be used for designing inhibitors with higher efficiencies against the the insect proteases. These findings may help develop new strategies for the management of Eurygaster integriceps Puton.
2. Methods
2.1. PEP sequences
The protein sequence of PEP, accession number ACI03586.2, was obtained from the National Centre for Biotechnology Information (NCBI) database to perform analysis. Sequences from GenBank, RefSeq, TPA, SwissProt, PIR, PRF, and PDB are available in the comprehensive NCBI Protein database (http://www.ncbi.nlm.nih.gov/protein). Numerous analyses and investigations were made feasible by the NCBI Protein database's retrieval of the particular protein sequence, ACI03586.2. Annotated translations of coding regions from GenBank, RefSeq, and TPA, along with records from SwissProt, PIR, PRF, and PDB, are all available in the Protein database. The wealth of protein sequences gathered from these varied sources forms the essential foundation for comprehending the structure and function of living things.
2.2. Homology modeling search
The PEP protein was analyzed using the NCBI BLAST tool (http://blast.ncbi.nlm.nih.gov/Blast.cgi), more especially the Basic Local Alignment Search Tool (BLAST). The PEP protein sequence was compared to a non-redundant protein database as part of this analysis. This BLAST analysis aimed to find similar regions and possible matches with other proteins in the database. Additionally, we looked into the PEP protein's possible conserved domains during our analysis. The parts of the protein sequence known as conserved domains, which show a high degree of similarity between various proteins, can reveal important information about the functional properties and evolutionary links of the protein. The NCBI BLAST tool was an invaluable resource for these investigations as it provided us with the capability we needed.
2.3. Template search
In order to find putative homologous structures, we used the NCBI BLAST tool and, more specifically, performed a PSI-BLAST analysis using the query protein sequence. The protein data bank (PDB) served as the target database for this analysis. By using PSI-BLAST, we sought to find proteins in the PDB that show a significant degree of sequence similarity to our query protein. This method allows for the identification of putative homologous structures, which can offer insights into the three-dimensional architecture and possible functions of the query protein. By combining the strength of PSI-BLAST with the PDB's richness, our objective was to shed light on the putative homologous structures of the query protein, which can help us understand its potential biological.
2.4. Sequence alignments
Four protein sequences of PEP, which showed several features including the highest total score > 1000, E-value = 0, and identity > 80% retrieved from the previous step were aligned for homology evaluation. For analysis of the validity of the same sequences, the amino acid sequence of Prolyl-endopeptidase was used for alignment against template sequences based on an already conducted template search. All alignment generations were performed using the CLC 6.1 software. The BLOSUM substitution matrix was selected with values of 10, and 0.1, respectively for gap penalty gap extension penalty.
2.5. Topology assessment
We used two online databases, TOPCONS and TMHMM, to identify the protein's hydrophobic transmembrane region. TOPCONS is a consensus prediction tool specifically made for membrane protein topology and signal peptide analysis. It may be accessed at https://topcons.cbr.su.se. When the protein's amino acid sequence is entered into TOPCON in FASTA format, it produces predictions about the membrane topology and finds putative signal peptides. We also used the TMHMM database, which is accessible at https://services.healthtech.dtu.dk/service.php, in addition to TOPCONS.TMHMM-2.0. Transmembrane Hidden Markov Model, or TMHMM for short, is a popular technique for identifying transmembrane helices in protein sequences. TMHMM uses a statistical model to detect putative transmembrane regions and determine how hydrophobic they are based on the protein sequence.
2.6. Secondary structure prediction
Our approach of choice for predicting protein secondary structures was the self-optimized prediction method (SOPM). We used SOPM (http://npsa-pbil.ibcp.fr/cgi-bin/npsa_automat.pl page = npsa_sopma.html ) to anticipate secondary structural components including alpha helices, beta strands, and coils by analyzing the protein sequence. Using a combination of statistical models and algorithms, SOPM is a computational technique that generates precise predictions based on the amino acid sequence.
2.7. PEP 3D structure prediction
Three approaches were used for modeling, which included SWISS-MODEL (available at https://swissmodel.expasy.org), (PS)2 at http://ps.life.nctu.edu.tw/index.php, Phyre2 (accessible via the online address: http://www.sbg.bio.ic.ac.uk/phyre2/html/page.cgiid=index, and I-TASSER (Iterative Threading ASSEmbly Refinement) (available at https://zhanggroup.org/ I TASSER). The final 3D structure was built by the modeling package MODELLER.
2.8. Models evaluations
To evaluate the models, we utilized QMEAN (Qualitative Model Energy Analysis), a widely used tool for assessing the quality of protein structures. QMEAN provides valuable insights into the major geometrical features of protein structures and helps to estimate their overall reliability and accuracy. The QMEAN analysis was performed using the online platform available at http://swissmodel.expasy.org/qmean/cgi/index.cgi, which offers a comprehensive set of evaluation parameters and scoring functions (21).
2.9. Model refinement
ModRefiner, an advanced tool accessible at http://zhanglab.ccmb.med.umich.edu/ModRefiner/, played a pivotal role in the refinement process of our models. This software specializes in constructing and refining protein structures starting from Cα traces using atomic-level energy minimization. By utilizing ModRefiner, we were able to enhance the accuracy and physical realism of our initial models. The refinement process involved iteratively adjusting the positions of both backbone and side-chain atoms, allowing for comprehensive flexibility during the simulations. This flexible conformational search was guided by a combination of physics-based and knowledge-based force fields, enabling the models to converge toward their native state (22).
2.10. Structure Alignment
To identify the best structural alignment and compare the predicted 3D structure of the protein with a template exhibiting favorable features, we employed the Dali web-based tool available at http://ekhidna.biocenter.helsinki.fi/dali_server/. Firstly, we selected the best-predicted 3D structure generated for the protein in our study. This structure was chosen based on its overall quality, reliability, and adherence to known structural principles. Subsequently, we utilized the Dali web-based tool to perform structural alignment between the selected 3D structure and the template with the most desirable characteristics. The template, identified by its PDB code as 3qlB, exhibited satisfactory features that made it a suitable reference for comparison.
2.11. Ligand binding site prediction
In our study, we employed COFACTOR, a widely used tool available at http://zhanglab.ccmb.med.umich.edu/COFACTOR/, for the prediction of ligand binding sites. COFACTOR is a comprehensive method that integrates structure, sequence, and protein-protein interaction information to annotate the biological roles of proteins. Using COFACTOR, we aimed to identify and characterize the specific regions within the protein structure that are involved in binding to ligands. Ligand binding sites play crucial roles in protein function, as they are responsible for interactions with small molecules, ions, or other proteins, often leading to important biological activities.
2.12 Interface residues and surface pockets prediction
we utilized two important tools for the analysis of protein structures. The first tool, InterProSurf, available at http://curie.utmb.edu/prosurf.html, was employed to identify interface residues in protein complexes. InterProSurf is a powerful server that specializes in predicting interacting sites on protein surfaces. By analyzing protein complex structures deposited in the Protein Data Bank (PDB), InterProSurf enables the identification of residues involved in protein-protein interactions. This information is valuable for understanding the functional and structural aspects of protein complexes. The second tool we utilized is GHECOM (Grid-based HECOMi finder), accessible at https://pdbj.org/ghecom/. GHECOM is a program designed to identify multi-scale pockets on protein surfaces using mathematical morphology. It employs a grid-based approach to detect these pockets, which can be potential binding sites for ligands or other molecules. By analyzing the protein surface at different scales, GHECOM allows the identification of various types of pockets, including small, medium, and large ones. This information provides insights into the structural characteristics and potential functional roles of these pockets.
2.13. Identification of residues with putative crucial structure and function
The consurf program (http://consurf.tau.ac.il/) was used for the identification of functional residues of the nanobody structure. The software parameters were set as PSI-BLAST for five iterations against the Uniprot database with an E-value of 0.01 and maximum likelihood (ML) to calculate the amino acid conservation score.
2.14. protein contacts, accessibility, and residue volume determination
For achievement of information like the protein packing, residue volume, contacts, accessibility, and topological genus calculations, VLDPws based on the Laguerre diagram (https://www.dsimb.inserm.fr/dsimb_tools/vldp/index.php) is a routine tool, which was utilized in this study.
3. Results
3.1. BLAST results
A
The BLAST search conducted using the query sequence yielded several significant hits, indicating sequence similarities with various organisms. Among the hits, notable matches included [Eurygaster integriceps], [Halyomorpha halys], [Apolygus lucorum], [Cimex lectularius], [Zootermopsis nevadensis], and several others. This finding suggests that the query sequence shares homology with these organisms, indicating potential evolutionary relationships or functional similarities. The identified hits represent species from diverse taxonomic groups, suggesting that the sequence may have conserved regions that are important for its biological function. Further analysis of the query sequence revealed the presence of putative conserved domains associated with the Prolyl oligopeptidase PreP. Conserved domains are specific regions within a protein sequence that are highly conserved across different species and are often associated with particular functions or structural features. In this case, the identified conserved domains are characteristic of the Prolyl oligopeptidase PreP protein.
3.2. Template selection
Table 1 represents the first 10 hits having the highest BLAST scores on the query sequence against the PDB database. The first hit (Accession: 6JCI_A, the Crystal structure of PEP from Haliotis discus hannai with SUAM-14746 [Haliotis discus hannai], Max score: 734, Query coverage: 99%, Max ident: 50.35%) showed the best score and so was selected as a template.
Table 1
The first 10 hits with the highest BLAST scores on the query sequence against Protein Data Bank (PDB).
Description
Scientific Name
Max Score
Total Score
Query Cover
E value
Per. Ident
Acc. Len
Accession
Crystal structure of Prolyl Endopeptidase from Haliotis discus hannai with SUAM-14746 [Haliotis discus hannai]
Haliotis discus hannai
734
734
99%
0.0
50.35%
758
6JCI_A
Prolyl Oligopeptidase From Porcine Brain [Sus scrofa]
Sus scrofa
721
721
99%
0.0
50.21%
710
1H2W_A
Prolyl Oligopeptidase From Porcine Brain, T597c Mutant [Sus scrofa]
Sus scrofa
721
721
99%
0.0
50.21%
710
1VZ3_A
Prolyl oligopeptidase from porcine brain, D641N mutant with bound peptide ligand SUC-GLY-PRO [Sus scrofa]
Sus scrofa
721
721
99%
0.0
50.21%
710
1O6G_A
POP from porcine brain, mutant, complexed with inhibitor [Sus scrofa]
Sus scrofa
721
721
99%
0.0
50.07%
710
1E8M_A
POP from porcine brain, Y473F mutant [Sus scrofa]
Sus scrofa
720
720
99%
0.0
50.07%
710
1H2X_A
POP from porcine brain, D641A mutant with bound peptide ligand SUC-GLY-PRO [Sus scrofa]
Sus scrofa
718
718
99%
0.0
50.07%
710
1O6F_A
Prolyl Oligopeptidase From Porcine Brain, H680a Mutant [Sus scrofa]
Sus scrofa
718
718
99%
0.0
50.07%
710
4AX4_A
Chain A, Prolyl endopeptidase [Homo sapiens]
Homo sapiens
717
717
99%
0.0
49.86%
709
3DDU_A
POP from porcine brain, mutant [Sus scrofa]
Sus scrofa
715
715
99%
0.0
49.93%
710
1E5T_A
3.3. Alignments
When analyzing the sequence homology between PEP and the selected template, a schematic illustration in Fig. 1 was used to visually represent the comparison. The results of this comparison revealed that PEP and the template share a sequence identity of 50.35%. Sequence homology refers to the degree of similarity or identity between two protein sequences. In this case, PEP and the template were aligned and compared against each other, examining the matching residues and their positions. The sequence identity of 50.35% indicates that approximately half of the amino acid residues in PEP align with equivalent positions in the template sequence.
Fig. 1
Comparing the sequences' homology. A schematic illustration of homology between the protein sequence and the selected template.
Click here to Correct
3.4. Topology modeling
To gain insights into the structural organization of PEP, a topology model was constructed. The topology model provides information about the arrangement of different regions within the protein, including the presence or absence of transmembrane segments. In this case, the topology model for PEP indicated the absence of any transmembrane regions within the protein (Fig. 2). The absence of transmembrane regions suggests that PEP is likely a soluble protein, rather than being embedded within cell membranes. Soluble proteins are typically found in the cytoplasm, nucleus, or other cellular compartments where they perform various enzymatic or regulatory functions.
Fig. 2
Topology model of protein. No transmembrane region for the PEP was predicted.
Click here to Correct
3.5. Secondary structure
the evaluation of the protein secondary structure revealed the presence of multiple structural elements, including alpha helices, extended strands, beta-turns, and random coils. The proportions of these secondary structure components were determined, with alpha helix, extended strand, beta-turn, and random coil comprising approximately 24.75%, 26.44%, 8.02%, and 40.79% of the structure, respectively. The graphical representation in Fig. 3 visually depicts the distribution of these secondary structure elements, providing a comprehensive overview of the protein's structural composition.
Fig. 3
Secondary structure prediction. Alpha helix (24.75%), extended strand (26.44%), beta-turn (8.02%), and random coil (40.79%).
Click here to Correct
3.6. 3D modeling
Swiss model and ps2v2 recruited for homology modeling created two models. Phyre2 predicted one 3D model and I-TASSER suggested 5 models. I-TASSER simulations created a high number ensemble of structural conformations termed decoys. To screen the estimated models, I-TASSER employs the SPICKER program for clustering all the decoys based on the similarity of pair-wise structures and suggests several models up to five numbers corresponding to the five largest structure clusters. The confidence of models was quantitatively measured by calculation of the C-score according to the significance of threading template alignments and the convergence parameters of the structure assembly simulations. The normal range of the C-score is [-5, 2], where a higher value indicates a model with higher confidence and vice-versa. Moreover, the TM-score and RMSD (root-mean-square deviation) were calculated based on the C-score, and the protein length is determined by the association between these qualities. The model with the higher C-score (C-score = 1.42, Estimated TM-score = 0.91 ± 0.06, Estimated RMSD = 5.0 ± 3.3Å) was selected for further analysis.
3.7. Evaluation of models
The QMEAN server was employed to estimate the 3D models of the protein. Among the predicted models generated by different methods, including phyre2, Ps2v2, and I-Tasser, the Swiss Model 3D structure obtained the highest QMEAN score, indicating its superior quality. The graphical representation in Fig. 4 and the data in Table 2 serve to visually and quantitatively demonstrate the comparison of QMEAN scores for the different predicted models. The selection of the Swiss Model 3D structure based on its higher QMEAN score suggests that it is the most reliable and accurate model for further analysis.
Fig. 4
Local quality estimate of 3D models predicted by the Swiss model, Ps2v2, phyre2, and I TASSER (up to down, respectively). The model with the higher C-score (C-score = 1.42, Estimated TM-score = 0.91 ± 0.06, Estimated RMSD = 5.0 ± 3.3Å) was selected for further analysis.
Click here to Correct
Table 2
QMEANDisCo score predicted for the Swiss model, phyre2, Ps2v2, and I tasser models.
Server
swiss model
Ps2v2
phyre2
I tasser
QMEANDisCo score
0.82
0.78
0.79
0.75
3.8. Model refinement
The 3D structure of the protein initially existed as the primary model, which served as an initial approximation. Through a refinement process, the final model was obtained, representing an improved version with enhanced structural quality. The graphical representation in Fig. 5 visually displays the primary and final models, allowing for a direct comparison of the structural improvements achieved through refinement.
Fig. 5
3D structure of the initial and the final model after refinement.
Click here to Correct
Click here to Correct
3.9. Structures alignments
No significant difference was found between the structures built for the template and the 2D and 3D structures of the query protein. The comparison of their tertiary structures revealed matches in alignment, indicating a close similarity in the overall folding patterns and spatial arrangements of the secondary structure elements. The graphical representation in Fig. 6 visually demonstrates the alignment of the tertiary structures, further supporting the observed similarities between the template and the query protein.
Fig. 6
Dali 3D structure alignment. Structure alignment between Prolyl-endopeptidase (green) and 6JCI_A (orange) from lateral and top views, respectively. Ligand appears in the space-filling model in red color.
Click here to Correct
3.10. Ligand binding site predictions
The COFACTOR algorithm was utilized to predict the ligand binding sites within the protein. The analysis identified several conserved residues, including positions 173, 472, 475, 553, 554, 599, 644, 645, and 681, that are likely involved in ligand binding. The graphical representation in Fig. 7 visually illustrates the predicted ligand binding sites and highlights the positions of the conserved residues within these sites. This information provides valuable insights into the protein's functional characteristics and potential for ligand interactions.
Fig. 7
COFACTOR Ligand binding site prediction. Involvement of conserved residues, especially 173, 472, 475, 553, 554, 599, 644, 645, and 681 in ligand binding sites.
Click here to Correct
3.11. Interface residues prediction
The Interprosurf tool was utilized to find interface residues within the protein complex structure. Residues 566, 567, 568, 569, and 570 were predicted as interface residues and are depicted in Fig. 8. Additionally, the GHECOM server identified five pockets on the PEP surfaces of the protein complex, and these pockets are shown in Fig. 9. The information obtained from these analyses contributes to a better understanding of the protein complex's interactions, potential binding sites, and functional characteristics, facilitating further investigations in the field of molecular biology and drug discovery.
Fig. 8
InterProSurf result. Automatic Patch Analysis predicts the following residues: 566, 567, 568, 569, 570
Click here to Correct
Fig. 9
GHECOM results showing graph residue-based pocketness and Jmol view of the pocket structure. Top: graph residue-based pocketness. The bar height shows the value of pocketness [%] for each residue. The color of the pocketness bar indicates the cluster number of pockets (red: cluster 1, blue: cluster 2, green: cluster 3, yellow: cluster 4, cyan: cluster 5). Below: Jmol view of pocket structure based on cluster color (left) and pocketness color(right).
Click here to Correct
3.12. Identification of residues with putative crucial function and structure
Consurf analysis was conducted to identify functional residues within the nanobody. Both the sequence and structure of the nanobody were analyzed to assess the degree of conservation and identify functionally important positions. The results of this analysis are presented in Fig. 10, offering a visual representation of the functional residues within the nanobody. These findings contribute to a better understanding of the nanobody's functional characteristics and can guide further experimental investigations and applications in various fields, such as antibody engineering and therapeutics.
Fig. 10
Identification of functionally and structurally important residues predicted by the ConSurf server. Annotated functional residues on structure in the twilight zone.
Click here to Correct
3.13. protein contacts, accessibility, and residue volume determination
The VLDP server was used to generate a contact map depicting the interactions between residues within the protein. The contact map in Fig. 11 provides a visual representation of these interactions. Additionally, the quantification of the contacts, represented by a blue scale legend, offers a numerical measure of the common area between the atoms of each residue. Furthermore, the VLDP server computed additional information about the contacts, including exposed surface area, total surface area, neighboring residues, sequence positions, and amino acid types. These details enhance the understanding of the contact patterns and contribute to the characterization of the protein's structure and functionality.
Fig. 11
Contacts. The VLDP server provided the contact. A blue scale legend gives the number of contacts.
Click here to Correct
The Laguerre diagram, utilizing finely tuned parameters, enables the determination of accurate volume values for each residue within the protein. Table 3 presents the volume measurements for individual residues, providing insights into their spatial extent. The interactive nature of the volume data allows for dynamic exploration and analysis. These volume measurements contribute to a comprehensive understanding of the protein's structure and can be utilized to investigate its functional implications and molecular interactions.
Table 3
A part of the Volume & Area table. MOL: Name of Molecule; NUMMOL: Sequence of MOL; AREA: Total area of MOL; VOLUME: Volume of MOL; LOCUS: 0 = homocontacts (AA-AA, or HOH-HOH); 1 = heterocontacts (AA-WAT); 2 = contact with vertices of boxes.
MOL
NUMMOL
AREA
VOLUME
LOCUS
MET
1
353.29708
169.16622
1
LYS
2
384.42640
170.84407
1
LYS
3
350.70454
152.08738
1
PHE
4
483.13079
210.98841
1
GLN
5
332.71736
136.34521
1
TYR
6
497.14338
204.98338
1
PRO
7
268.94889
118.74531
1
GLU
8
332.71428
137.45408
1
ALA
9
204.84493
92.472418
1
ARG
10
407.77550
168.64672
1
The LOCUS column provides information about the nature of residue contacts, distinguishing between buried, solvent-exposed, and boundary residues. This information aids in understanding the accessibility and potential functional roles of residues within the protein. The PIA analysis, depicted in Fig. 12 and presented in Table 4, provides a detailed characterization of the residue contacts, including interaction types and strengths. Integrating these findings enhances the understanding of the protein's structure, interactions, and potential functional mechanisms.
Fig. 12
Accessibility (PIA). An extra column named LOCUS provides information about the residue contacts.
Click here to Correct
Table 4
PIA table. A part of the Accessibility (PIA) table. AA: Name of the Amino Acid; resseq: residue sequence number; PIA: Polyhedra Interface Area = ESA/TSA; TSA: Total Surface Area of AA; ESA: Surface Area Exposed to Solvent.
# AA
Resseq
0%
M
1
63%
K
2
46.7%
K
3
77.2%
F
4
22.4%
Q
5
71.9%
Y
6
13.7%
P
7
42.4%
E
8
78.9%
4. Discussion
The increasing human population, and nutrition crisis in addition to the inefficiency and adverse effects of available pesticides have warned of the urgent need for new and efficient pesticides (13, 23). Insect pests hurt nearly one-third of agricultural crops and forestry production all over the World (24). Sunn Pest is one of wheat's most harmful insects and causes remarkable hurt to this worth crop annually (8). This pest feeds on different parts of the developing wheat and barley plant, injects PEP, and eventually degrades the wheat gluten (2).
PEP is one of the important enzymes commonly found in the salivary gland of this insect (25). This PEP is a serine protease of the S9A family with the ability to digest wheat gluten proteins to smaller particles (26, 27). PEP cleaves at internal proline residues within gluten and gluten-like molecules found in wheat, rye, and barley (2, 5). Almost 80% of wheat protein is gluten, so PEP can play an important role in feeding Sunn pests (2). Darkoh et al. showed that PEP exhibited a Km value of 65.3 ± 1.8 µM against the substrate GlyPro-pNA at pH 8 in ethanolamine buffer (5).
Gluten consists of gliadins and glutenin peptides with proline and glutamine-rich sequences (2, 5). This disordered elastomeric protein is responsible for dough's remarkable viscoelastic properties. Viscosity and elasticity of gluten are crucial in the food industry (28).
Several studies on other digestive enzymes, including α-amylase, have been performed, which are discussed below. It was showed that Cs’ natural amylase inhibitors have inhibitory activity moderate to higher against Tribolium Castanamylase α-amylase (2932). The α-amylase inhibitors are found in microbes, animals, and numerous plant seeds and tubers, especially cereals and legumes (33). Pereira et al. unveiled the crystal structure of α-amylase extracted from yellow mealworm in a complex with an Amaranth inhibitor. This unique inhibitor might provide a valuable insecticide for protecting agricultural products (34). Considering that α-amylase inhibitor is present in cereals but did not control Sunn Pests, it can be mentioned that the alpha-amylase inhibitor may not be a suitable option to control this insect.
Saadati et al. showed the effects of several proteinase inhibitors, including tosyl-L-lysine chloromethyl ketone (TLCK), and N- tosyl-L-phenylalanine chloromethyl Ketone (TPCK) on Sunn Pest gut serine proteinase. Blocking proteases leads to amino acid deficiency and reduced development, growth, and fecundity. This study showed that to achieve the desired results in managing Sunn Pests, it should combine two inhibitors (26).
Oppert et al. showed that combining potato cysteine proteinase inhibitors with soybean trypsin inhibitors is more efficient than either inhibitor alone in inhibiting the development and survival of T. castaneum larvae. These insecticides are indicated as promising agents for developing transgenic seeds that are resistant to plant pests (27).
Another study showed that several inhibitors targeting insect cysteine and serine proteinases have a synergistic inhibitory impact on T. Castaneum mid-gut proteolytic activity and hamper to hurt to stored crops (35).
One approach to gaining accurate knowledge is to use bioinformatics software that can lead to the identification of more potential targets and the development of insecticides that are selective and specific. In silico approaches can accelerate the discovery of specific and novel pesticides that will result in an impressive increase in the number of discovered pesticides with higher efficiency (23).
The necessity of PEP in feeding and survival of Sunn pest in winter provided us the rationale to explore its structure as an inhibitor target candidate. Bioinformatics approaches were used to achieve such a target.
Our BLAST results revealed that PEP was found in several pests, including Halyomorpha halys, Apolygus lucorum, Cimex lectularius, and Zootermopsis nevadensis. PEP inhibitors cross-react with a various pest for high identity reasons. To achieve the best template, the PEP sequence was used as a query for a BLAST search against the PDB. The PEP from Haliotis discus hannai (SUAM-14746) with the highest concession was chosen as a template. A hit showing the best whole score can be regarded as the most reputable template.
The first step in structure prediction is identifying a relation between the target sequence and possible templates (36). The proportion of secondary structure in the PEP is alpha helix (24.75%), extended strand (26.44%), beta-turn (8.02%), and random coil (40.79%). Therefore, the random coil is the major region in this protein's secondary structure.
The precision of forecast homology modeling is determined by the degree of sequence resemblance. Homology modeling could be the best if the structure template with query protein finds an identity of > 50% (37). Since the similarity between the query and its template sequence was 50.35% in this study, we suggested that homology modeling can be more potent than the other methods.
ModRefiner uses an algorithm for high-resolution refining of protein structure, which can significantly improve local structures' physical quality (22, 36). In the topology model presented here, this protein doesn’t have any transmembrane region.
Surface binding pockets are an attractive target for drug and inhibitor design against the pathogen, five pockets on protein surfaces were obtained using GHECOM server mathematical morphology. Auto Patch Analysis predicted residues 566, 567, 568, 569, and 570.
COFACTOR software is used for finding ligand binding sites indicating the involvement of conserved residues, especially 173, 472, 475, 553, 554, 599, 644, 645, and 681 in the ligand binding site.
Collectively, our study provides information on the structure of this vital enzyme retrieved using several bioinformatics approaches. The structure introduced here facilitates further research to elucidate thoroughgoing PEP as an inhibitor target for managing Sunn Pests. Due to the resistance of pests to common insecticides, as well as obtaining more favorable results with the combination of different insecticides, the need to identify and design new insecticides is urgent and unmet. Accurate and comprehensive identification of the structure of insect digestive enzymes is an essential step in designing new inhibitors with high efficiency for insect management, including Sunn Pests (2).
5. Conclusion
The increasing human population, along with a nutrition crisis and limitations of existing pesticides, highlights the urgent need for new and effective pesticides. Insects, particularly the destructive Sunn Pest, pose a significant threat to global agricultural and forestry production. The Sunn Pest injects PEP, a serine protease, which degrades the gluten in wheat crops. This study explores the potential of PEP as a target for developing inhibitors to control Sunn Pests. Bioinformatics tools were used to analyze PEP and generate a model that identified key binding sites and conserved residues. This research provides valuable insights into the structure of PEP and serves as a foundation for further studies on developing effective insecticides. Given the resistance of pests to existing chemicals, the identification and design of new insecticides are crucial for effective pest management. Understanding insect digestive enzyme structures is a vital step in designing potent inhibitors for controlling Sunn Pests.
A
Author Contribution
Conceptualization, Effat Noori; methodology and software, Fatemeh Sefid and Effat Noori; validation, Fatemeh Sefid; formal analysis, Effat Noori; resources, Effat Noori; data curation, Sajad Najafi and Ali Ahmadizad Firouzjaei; writing—original draft preparation, Effat Noori and Ali Ahmadizad Firouzjaei; writing—review and editing, Sajad Najafi; supervision, Mojgan Bandehpour and Bahram Kazemi; project administration, Bahram Kazemi; funding acquisition, Bahram Kazemi. All authors have read and agreed to the published version of the manuscript.
A
Funding:
This paper was taken from Ms. Effat Noori’s PhD thesis and financially supported by the Cellular and Molecular Biology Research Center of Shahid Beheshti University of Medical Sciences of Iran [grant no. 14160].
A
Data Availability
All relevant data are included in the article.
Acknowledgments:
We thank the services provided by the biotechnology department of the School of Advanced Technologies in Medicine, Shahid Beheshti University of Medical Sciences.
Conflicts of interest:
The authors declare that there is no conflict of interest.
References
1.
Konarev A, Dolgikh V, Senderskiy I, Konarev A, Kapustkina A, Lovegrove A (2019) Characterisation of proteolytic enzymes of Eurygaster integriceps Put.(Sunn bug), a major pest of cereals. J Asia Pac Entomol 22(1):379–385
2.
Darkoh C, El-Bouhssini M, Baum M, Clack B (2010) Characterization of a prolyl endoprotease from Eurygaster integriceps Puton (Sunn pest) infested wheat. Arch Insect Biochem Physiol 74(3):163–178
3.
Zibaee A, Bandani A (2010) Effects of Artemisia annua L.(Asteracea) on the digestive enzymatic profiles and the cellular immune reactions of the Sunn pest, Eurygaster integriceps (Heteroptera: Scutellaridae), against Beauveria bassiana. Bull Entomol Res 100(2):185–196
4.
Alizadeh M, Sheikhi-Garjan A, Ma’mani L, Hosseini Salekdeh G, Bandehagh A (2022) Ethology of Sunn-pest oviposition in interaction with deltamethrin loaded on mesoporous silica nanoparticles as a nanopesticide. Chem Biol Technol Agric 9(1):1–13
5.
Yandamuri RC, Gautam R, Darkoh C, Dareddy V, El-Bouhssini M, Clack BA (2014) Cloning, expression, sequence analysis and homology modeling of the prolyl endoprotease from Eurygaster integriceps Puton. Insects 5(4):762–782
6.
Genc H, Genc L, Turhan H, Smith S, Nation J (2008) Vegetation indices as indicators of damage by the sunn pest (Hemiptera: Scutelleridae) to field grown wheat. Afr J Biotechnol. ;7(2)
7.
Davari A, Parker BL (2018) A review of research on Sunn Pest {Eurygaster integriceps Puton (Hemiptera: Scutelleridae)} management published 2004–2016. J Asia Pac Entomol 21(1):352–360
8.
Hosseininaveh V, Bandani A, Hosseininaveh F, Cohen A (2009) Digestive proteolytic activity in the Sunn pest, Eurygaster integriceps. J Insect Sci. ;9(1)
9.
Cantón PE, Bonning BC (2019) Proteases and nucleases across midgut tissues of Nezara viridula (Hemiptera: Pentatomidae) display distinct activity profiles that are conserved through life stages. J Insect Physiol 119:103965
10.
Iranipour S, Bonab ZN, Michaud JP (2010) Thermal requirements of Trissolcus grandis (Hymenoptera: Scelionidae), an egg parasitoid of sunn pest. Eur J Entomol 107(1):47
11.
Iranipour S, Pakdel AK, Radjabi G, Michaud J (2011) Life tables for sunn pest, Eurygaster integriceps (Heteroptera: Scutelleridae) in Northern Iran. Bull Entomol Res 101(1):33–44
12.
Noori E, Bandehpour M, Zali MR, Kazemi B (2023) In vitro Gluten Degradation Using Recombinant Eurygaster Integriceps Prolyl Endoprotease: Implications for Celiac Disease. Iran J Biotechnol 21(3):24–32
13.
Ahmadizad Firozjaei SA, Latifi AM, Khodi S, Abolmaali S, Choopani A (2015) A review on biodegradation of toxic organophosphate compounds. J Appl Biotechnol Rep 2(2):215–224
14.
Zharkov D, Nizamutdinov T, Dubovikoff D, Abakumov E, Pospelova A (2023) Navigating Agricultural Expansion in Harsh Conditions in Russia: Balancing Development with Insect Protection in the Era of Pesticides. Insects 14(6):557
15.
Critchley BR (1998) Literature review of sunn pest Eurygaster integriceps Put.(Hemiptera, Scutelleridae). Crop Prot 17(4):271–287
16.
Zhu F, Li XX, Yang SY, Chen YZ (2018) Clinical success of drug targets prospectively predicted by in silico study. Trends Pharmacol Sci 39(3):229–231
17.
Yousafi Q, Sarfaraz A, Khan MS, Saleem S, Shahzad U, Khan AA et al (2021) In silico annotation of unreviewed acetylcholinesterase (AChE) in some lepidopteran insect pest species reveals the causes of insecticide resistance. Saudi J Biol Sci 28(4):2197–2209
18.
Petrey D, Honig B (2005) Protein structure prediction: inroads to biology. Mol Cell 20(6):811–819
19.
Saeidnia S, Manayi A, Abdollahi M (2015) From in vitro experiments to in vivo and clinical studies; pros and cons. Curr Drug Discov Technol 12(4):218–224
20.
Khazaei-Poul Y, Farhadi S, Ghani S, Ahmadizad SA, Ranjbari J (2021) Monocyclic Peptides: Types, Synthesis and Applications. Curr Pharm Biotechnol 22(1):123–135 10.2174/1573412916666200120155104. PubMed PMID: 31987019
21.
Benkert P, Tosatto SC, Schomburg D (2008) QMEAN: A comprehensive scoring function for model quality assessment. Proteins Struct Funct Bioinform 71(1):261–277
22.
Xu D, Zhang Y (2011) Improving the physical realism and structural accuracy of protein models by a two-step atomic-level energy minimization. Biophys J 101(10):2525–2534
23.
Birgül Iyison N, Shahraki A, Kahveci K, Düzgün MB, Gün G (2021) Are insect GPCRs ideal next-generation pesticides: opportunities and challenges. FEBS J 288(8):2727–2745
24.
Guo D, Luo J, Zhou Y, Xiao H, He K, Yin C et al (2017) ACE: an efficient and sensitive tool to detect insecticide resistance-associated mutations in insect acetylcholinesterase from RNA-Seq data. BMC Bioinformatics 18(1):1–9
25.
Mehrabadi M, Bandani AR, Allahyari M, Serrão JE (2012) The Sunn pest, Eurygaster integriceps Puton (Hemiptera: Scutelleridae) digestive tract: Histology, ultrastructure and its physiological significance. Micron 43(5):631–637
26.
Saadati F, Bandani AR (2011) Effects of serine protease inhibitors on growth and development and digestive serine proteinases of the Sunn pest, Eurygaster integriceps. J Insect Sci 11(1):72
27.
Oppert B, Morgan T, Hartzer K, Lenarcic B, Galesa K, Brzin J et al (2003) Effects of proteinase inhibitors on digestive proteinases and growth of the red flour beetle, Tribolium castaneum (Herbst)(Coleoptera: Tenebrionidae). Comp Biochem Physiol C: Toxicol Pharmacol 134(4):481–490
28.
Dahesh M, Banc A, Duri A, Morel M-H, Ramos L (2014) Polymeric assembly of gluten proteins in an aqueous ethanol solvent. J Phys Chem B 118(38):11065–11076
29.
Pandey A, Yadav R, Sanyal I (2022) Evaluating the pesticidal impact of plant protease inhibitors: lethal weaponry in the co-evolutionary battle. Pest Manag Sci 78(3):855–868
30.
de Lira Pimentel CS, Albuquerque BNL, da Rocha SKL, da Silva AS, da Silva ABV, Bellon R et al (2022) Insecticidal activity of the essential oil of Piper corcovadensis leaves and its major compound (1-butyl‐3, 4‐methylenedioxybenzene) against the maize weevil, Sitophilus zeamais. Pest Manag Sci 78(3):1008–1017
31.
Wang B, Yang Y, Liu M, Yang L, Stanley DW, Fang Q et al (2019) A digestive tract expressing α-amylase influences the adult lifespan of Pteromalus puparum revealed through RNAi and rescue analyses. Pest Manag Sci 75(12):3346–3355
32.
Bruno D, Bonelli M, Cadamuro AG, Reguzzoni M, Grimaldi A, Casartelli M et al (2019) The digestive system of the adult Hermetia illucens (Diptera: Stratiomyidae): morphological features and functional properties. Cell Tissue Res 378:221–238
33.
Sivakumar S, Mohan M, Franco O, Thayumanavan B (2006) Inhibition of insect pest α-amylases by little and finger millet inhibitors. Pestic Biochem Physiol 85(3):155–160
34.
Pereira PJB, Lozanov V, Patthy A, Huber R, Bode W, Pongor S et al (1999) Specific inhibition of insect α-amylases: yellow meal worm α-amylase in complex with the Amaranth α-amylase inhibitor at 2.0 Å resolution. Structure 7(9):1079–1088
35.
Oppert B, Morgan T, Hartzer K, Kramer K (2005) Compensatory proteolytic responses to dietary proteinase inhibitors in the red flour beetle, Tribolium castaneum (Coleoptera: Tenebrionidae). Comp Biochem Physiol C: Toxicol Pharmacol 140(1):53–58
36.
Sefid F, Rasooli I, Jahangiri A (2013) In silico determination and validation of baumannii acinetobactin utilization a structure and ligand binding site. BioMed research international. ;2013
37.
Floudas C, Fung H, McAllister S, Mönnigmann M, Rajgaria R (2006) Advances in protein structure prediction and de novo protein design: A review. Chem Eng Sci 61(3):966–988
Total words in MS: 5571
Total words in Title: 19
Total words in Abstract: 150
Total Keyword count: 5
Total Images in MS: 12
Total Tables in MS: 4
Total Reference count: 37