The ScriptTiO2.Rmd file is seperated in three parts i.e. Part 1, Part 2 and Part 3 (described below). The necessary data files can be found in the data folder. The orginal data, retrieved from GEO, accession:GSE42069 and analysed with ArrayAnalysis, can be found in the original_data subfolder of the aforementioned data folder. The files used in the Data pre-processing part (described below) can be found in the data_pre_processing subfolder of the aforementioned data folder. The resulting files can be found in the output folder. The functions that are used in the ScriptTiO2.Rmd file are found in the functions folder.
Pre-processing of data files, which can be found in the original_data subfolder. This part gene identifiers so that the files include Ensemble identifiers, HGNC symbols and Entrez Gene identifiers. Moreover, it merges the separate files into one file named TiO2-dataset.txt and saves this file in the output folder. Furthermore it creates depict signififcantly differentially expressed genes in volcano plots for all conditions, using the EnhanvedVolcano package in R (version 3.6.1) [Blighe, Rana, and Lewis (2018)].
For the next three parts use the ScriptTiO2.Rmd file
The first part is used to find affected pathways. This part uses overrepresentation analysis (ORA) based on gene expression data. This method is used to identify significantly enriched pathways from the WikiPathways database. In this part the enricher function of the clusterProfiler package in R [PMID:22455463]. Almost all variables were kept standard except the Adjusted p-value cut-off and q-value cut-off were set at lower than 0.05 and minimal gene set and maximal gene set size were set to 10 and 300 respectively. As mentioned above, this part makes use of the pathway models from the WikiPathways database (www.wikipathways.org). It is an open platform for the curation of biological pathways and hosts a pathway database with custom graphical pathway editing tools [PMID:29136241]. For both this part and the second part, pathway-gene information was obtained in the Gene Matrix Transposed (GMT) file format from http://data.wikipathways.org/. The pathway dataset included the Curated Collection and the Reactome pathways.
The second part is used to find toxicity related pathways based on their genetic overlap with genes from the Gene Ontology (GO) terms. For this part the GO-terms were selected based on their relation to toxicologic processes i.e. responses in the human body. Genes were obtained from the Gene Ontology (GO) (AmiGO 2 version 2.5.12) categories “apoptotic process” (GO:0006915), “inflammatory response” (GO:0006954), “cellular response to DNA damage stimulus” (GO:0006974) and “response to oxidative stress” (GO:0006979) available at www.geneontology.org [PMID:10802651, 30395331]. The genes related to these four GO-terms are retrieved using the biomaRt package available in R (version 2.42.0) [PMID:19617889, 16082012]. ORA for comparison as an approach compared to genetich overlap was done similar as described above in part 1 and the WikiPathways database was also used as described above.
The third part combines the results of the previous two parts. This part uses the results of the second part to filter those of the first part. This approach leads to toxicity related pathways which are used for further biological analysis.
N.B. to run the workflow for the 1h time point the 1h repository can be used: