Wednesday, October 5, 2016

SysBio16 Asgn_14B_Class_14_Article_17_2016_10_11

Mass Spectrometry: Read Article 17  Cody R. Goodwin, Stacy D. Sherrod, Christina C. Marasco, Brian O. Bachmann, Nicole Schramm-Sapyta, John P. Wikswo, and John A. McLean. Phenotypic Mapping of Metabolic Profiles Using Self-Organizing Maps of High-Dimensional Mass Spectrometry Data. Anal.Chem. 86 (13):6563-6571, 2014.. This is an example of metabolomics. It also contains a picture of two principal components - be sure you understand it. Post a PCRC.

17 comments:

  1. 0. Knew:
    I knew the basis of SOMs and having a background from class on PCA was very helpful going through this paper. Familiar with using MS to extract the data for these types of experiments (although not very knowledgeable on actually performing MS.. methods section was helpful here).

    1. Learned:
    Learned other applications of using SOM strategies (intro.). Learned how a lot of the techniques we've discussed can be used collectively to describe a data set. Also basic knowledge on how MEDI heat maps are constructed and how to extract information from a MEDI heat maps to compare different groups.

    2. Pressing ?:
    I'm a little bit confused on how the S plot in Figure 3b is used and what extra information we get from it. Did the data in the S plot come from PCA?

    3. Presentation:
    "Thresholding and feature recognition software..for future applications" (pg. 6568) benefits of this feature and what the data would look like.

    4. Thoughts:
    I liked that this paper used an experimental data set to describe MEDI rather than just trying to explain how it worked without the sample data. Made it MUCH easier to follow and understand applicability.

    ReplyDelete
    Replies
    1. We will hopefully talk about the s-plot during class!

      Delete
  2. Chinowsky_TheorSysBio_PCRC_101116

    0. Knew

    We had previously discussed self-organizing maps in the context of other papers, and I’ve read previous papers on mass spec for other classes. Additionally, our discussions in class about PCA were very helpful in the context of this paper.

    1. Learned

    The introduction to this paper discusses use of various multivariate statistical techniques, including PCA and OPLS-DA, which can be used, but are inadequate when it comes to complex interactions seen in high dimensional data sets. This paper uses the concepts of SOM as a strategy to look at “complex, nonlinear statistical relationships” and uses them to convert the high dimensional data in simple relationships in the context of metabolomics. The authors use a SOM workflow plan to simplify and visualize data from a rat model of cocaine addiction, thus allowing the visualization of metabolic phenotype and feature patterns from this model study. Specifically, Molecular Expression Dynamics Investigator (MEDI) was the workflow used. After data was initially mass corrected, centroided, and filtered, various software was used to indicate regions of interest and node location of features.

    2. Pressing

    It seems as though the MEDI heat map mainly allowed for a quick way to determine the statistically significant regions of the data set, thus allowing researcher to follow up with further statistic tests, such as one-way ANOVA. Is this the main advantage of such maps?

    3. Presentation

    Mass Spec processing

    4. Thoughts

    I enjoying seeing the application of MEDI through the use of experimental data and the conclusions that followed from the use of the technique. This made the paper easier to follow, and also allowed for insight into the practical applications of MEDI

    ReplyDelete
  3. 0. KNEW:
    I was familiar with SOMs and MEDI heat maps, but this paper helped me learn more about these different analysis techniques. I was also familiar with MS, but learned more about the methodology and applications of the analysis. I knew about the rat cocaine addiction model, but had never actually read a paper regarding the experiment.

    1. LEARNED:
    I learned more about SOMs and heat maps as well as there applications to the rat cocaine addiction model. I learned specifics about the MEDI heat map and how it can be used to compare different groups. Other metabolomics platforms can be applied to the data besides MS, such as NMR, multidimensional MS, or dynamic MS, but the significance of MS is understood after reading article 16.

    2. PRESSING QUESTIONS:
    In the abstract it is mentioned that Gestalt comparisons were used to prioritize features of the heat maps. What are these comparisons? When I looked up the definition of Gestalt psychology I learned that it is a theory that "tries to understand the laws of our ability to acquire and maintain meaningful perceptions in an apparently chaotic world." How were these features identified and how can these be prioritized without using previous assumptions?

    3. PRESENTATION:
    Group discussion of figures to help with clarification.

    4. THOUGHTS:
    It was an interesting application of heat maps and SOMs, the experiment was well-known enough to capture my attention for most of the paper.

    ReplyDelete
  4. 0. KNEW
    I have a general idea regarding SOMs and heat maps.

    1. LEARNED
    Multivariate statistical techniques (PCA and orthogonal partial least-squares discriminant analysis) can identify statistical correlations in data but cannot elucidate complex interactions in high-dim data sets. SOM converts the complex nonlinear statistical relationships between high-dim data into simple geometric relationships.
    I have a better idea on the workflow of the analysis of metabolomics data using a self-organizing map algorithm.

    2. QUESTIONS
    What is the difference between naive biological samples and non-addicted biological samples? I believe naive means rats that have not experienced cocaine?
    How is the heat map in Figure 3c self-organized?
    Why is it so important that tests be unsupervised and data-driven?

    3. PRESENTATION
    Comparisons between PCA, heatmaps, SOMs, etc. if realistic.

    4. THOUGHTS
    An interesting paper that utilizes heat maps and SOMs. I've never heard of the cocaine-rat experiment so it was fun to read about.

    ReplyDelete
  5. Ben Terrones

    SysBio16 Asgn_14B_Class_14_Article_17_2016_10_11

    0. KNEW:
    I knew how the measurements described (genomic, transcriptomic, metabolomics, proteomics) are used to uncover biological interactions. From previous discussions I knew about PCA and SOM.

    1. LEARNED:
    This paper described an experiment where SOM’s were used to visualize metabolomics data; this workflow was termed MEDI. Mass spectrometry was used in this case to acquire the data, but other metabolomics platforms, such as NMR, could have been used. Most of the paper discussed the actual experimental procedure in great detail, which helped to clarify some procedural questions. After describing the procedure, the authors went into the statistical analysis and construction of the MEDI heat map. For the mass spectrometry based data analysis, GEDI software was used.

    2. PRESSING ?:
    I was confused about how once you have the heat map and you know which regions are up-regulated or down-regulated, how you find out what specific metabolites are in those region?

    3. PRESENTATION:
    A presentation on some of the variations of MS (ICP-MS, IRMS, UPLC-IM-MS, etc.)

    4. THOUGHTS:
    I liked this paper because it gave an in depth example of using self-organizing maps and it was an interesting experiment.

    ReplyDelete
    Replies
    1. The colors of the heat map refer to the relative intensity of each compound, not comparing the entire data set to a control. We can then look at up- and down-regulation by comparing one of the maps to another.

      Delete
  6. Assn 14B, Article 17: Phenotypic Mapping of Metabolic Profiles Using Self=Organizing Maps of High-Dimensional Mass Spectrometry Data

    0. KNEW
    From previous papers, I knew how SOMs and PCA often lead to clustering of relevant data and separation of other aspects of the data. From the previous paper, I knew how different techniques could be combined with MS to create a multidimensional analysis, in this case, UPLC-IM-MS. In other papers, we have talked about how Euclidean distances and eigenvectors are used to study the raw data. I knew how the orthogonal vector to the eigenvector was used as the secondary PCA axis.

    1. LEARNED
    First of all, I learned that rats can get addicted to cocaine and can teach themselves how to self-administer cocaine. I learned that the SOMs techniques can be applied to data other than gene expression. I learned how the UPLC-IM-MS throughput is done and then how the data is edited (mass-correction, centroided, deisotoped, etc) before the PCA and OPLS-DA is performed. I found it interested how different compounds could be clustered on the MEDI heat map (Fig 3c). I learned how each peak in the SOM could be matched with a known compound via the mass found in MS and fragmentation data. I was interested that the SOM sorts data based on "similarity in intensity profiles across samples". I found it interesting that the 1st PCA was able to distinguish the naïve from the exposed rats and then that the 2nd PCA was able to distinguish the addicted from the non-addicted. I'm interested in the OPLS-DA and its ability to included so much of the sample variation, but I'm still a bit confused as to how it works. It was interesting to find out what kinds of metabolites were upregulated in the addicted rats.

    2. MOST PRESSING QUESTIONS
    How does OPLS-DA work and how does it include so much of the variation in the samples? Why does it seem like PCA is used more if OPLS-DA includes so much more of the variation? (pg 6568)
    The SOM algorithm is described as sorting the data points "based on similarity in intensity profiles across samples" but why isn't all of one color clustered together? Why are there multiple smaller clusters than form? (pg 6567)
    What was the data that corresponded to the eigenvectors? Can we know what data described the largest variations in the samples? (pg 6568)

    3. PRESENTATION TOPICS
    What do loading contributions represent and how does in influence the first principle component? (pg 6568).
    Other ways besides PCA and OPLS-DA to describe data from heat maps?

    4. THOUGHTS
    I liked than this papers used actual data to show their proof of concept. The info about the rats and the addiction made the technical parts of data analysis easier to understand. I'm still confused as to how the SOM algorithm works and what the PCA axis actually represents. I'm also curious if these results would be significant if the experiment was repeated with a larger sample size.

    ReplyDelete
  7. 0 Knew
    I had read this paper when I first joined the Bachmann lab, so coming into this time, I was pretty familiar with MEDI as a tool for identifying interesting features.

    1 Learned
    I really liked how this paper compared and contrasted the different methods of identifying unique features - between SOMs, PCA, and S-plots. Since I first joined my lab, I thought SOMs were magical, but it was difficult to intuit what each clustered represented as compared to PCA. Figure 4, connecting loadings analysis to MEDI, was v. helpful in this regard.

    2 Pressing Questions
    How were the statistically signficiant regions defined in fig 3c? Why are "f" "d" and "b" different regions as opposed to one common region. Will a random data set (or permutations of this dataset) generate a similarly organized map?

    3 Presentation Topic
    What is meant by this sentence? " The S- plot graphs features based on group specificity or correlation (ordinate) and covariance (abscissa). I thought correlation was just normalized covariance? How do you pronounce abscissa?

    4 Thoughts
    Very cool paper, helped me understand how PCA and MEDI are connected. I agree with @Natalie, the example with rat cocaine addiction was very intuitive and interesting, even if it was pedagogical (as the authors note).

    ReplyDelete
    Replies
    1. Abscissa is pronounced 'ab' (like the body part) 'sis' (like the immediate family member) 'uh' (like what you say when you don't know what to say). I am not sure about the actually important question.

      Delete
  8. 0. KNEW

    I was familiar with mass spectrometry and ion mobility from the previous class. I was also introduced to clustering and heat maps from a previous paper involving data sorting in terms of hierarchical clustering and dendrograms. I had experience in rat cocaine addiction behavior analysis from my work in Cold Spring Harbor Laboratory.

    1. LEARNED

    The paper first introduced multivariable statistical analysis methods that are used to compact many-dimensional datasets resulting from MS analysis of metabolomics of a sample. Prior to the experimental section, the paper presents Molecular Expression Dynamics Investigator (MEDI) workflow. A one-way ANOVA was described as comparing the effect of cocaine use by referring to clusters of signal ion intensities for the different classes of rats studied.

    The first steps of the workflow in the experiment conducted in this paper is data acquisition and preprocessing. This consisted of the rat conditioning, sacrifice, and sample analysis, as well as
    using software to filter, detect, and aline peaks across numerous MS platforms. Next, the workflow moved towards feature organization and analysis. Here, a self organizing map algorithm detected features (“defined as any detected monoisotopic molecular species with a discrete retention time and mass-to-charge ratio”) for production of multi-dimensioned grid. The process iterates through a random feature to generate a heat map of feature neighborhoods. The workflow ends with cluster decryption—determining features’ contributions to specific neighborhoods—and feature identification through comparison with well-known databases.

    2. PRESSING ?

    Unclear to me how the authors deduce that from Figure 3c, one can deduce that there are neighborhoods that “further discriminate cocaine-addicted and cocaine non-addictive models”?

    3. PRESENTATION

    LC, IM, MS in depth; more on SOMs

    4. THOUGHTS

    Dense and informative paper…I liked the experimental section because it was familiar, but I also appreciated the big data analysis portion

    ReplyDelete
  9. Class 14. Article 17: Cody R. Goodwin et al. Phenotypic Mapping of Metabolic Profiles Using Self-Organizing Maps of High-Dimensional Mass Spectrometry Data.

    0. KNEW:

    I was familiar with mass spectrometry as a common and powerful technique when trying to figure out what particular compounds are in a given sample (and especially so given the other paper we had to read this week). Also from previous class discussions and the other paper, I was familiar with self-organizing maps and principal component analysis.

    1. LEARNED:

    I got lost in the details of this paper at first, but I think I get the idea now: the paper is really just an illustration of using mass spectrometry to get data, and self-organizing maps and principal component analysis to analyze it (or as the paper says, the so-called 'MEDI workflow'). The important thing is not that cocaine addiction was studied in rats, but that you can gather data in a certain way, and that self-organizing maps can group data in biologically relevant ways.

    Let me mention some details for completeness' sake. Nine rats were studied; six were given cocaine for some amount of time, and then weened off of it, and three were never given any at all. Some of the rats that were given cocaine became addicted, and some weren't (and this was measured by how frequently they provided themselves cocaine). The rats were sacrificed, and their blood was put through a mass spectrometer to see what compounds were present, and at what concentrations.

    So now you've got a bunch of concentration data for a bunch of rats. It turns out that if you do PCA on this data, the first principal component separates the rats into two groups: those exposed to cocaine, and those that weren't! That's interesting given that it seems to imply that cocaine exposure leaves some sort of 'observable' mark. The next principal component separated rats into whether they were addicted or not.

    Such a 'mark' may help us identify what drug addiction means on a molecular level, which should surely be applicable to treatment one day.

    There were some other complementary analyses done (orthogonal partial least-squares-discriminant analysis was mentioned, for example), but I think the mass spectrometry and PCA stuff was most important, and that was my takeaway.

    2. QUESTIONS:

    This was something of a toy example; here, the principal components did not represent things we did not already know. Are there similarly 'macroscopic' examples along similar lines, where PCA tells us something surprising?

    3. PRESENTATION:

    The details of the experiment are not that important for us, I think. I would like to see a presentation on this discuss in detail the PCA and self-organizing maps-based analysis of the data, with maybe an argument about how some other common analyses don't tell you quite the same thing, or maybe they don't tell you the samr thing so easily (if that's true, anyway). Put another way, I want to see this workflow motivated. Is it really good at something others aren't?

    4. THOUGHTS:

    I hope I am right about what's important and not important. I have learned to 'read' biology papers by not reading them---or at least all of them. I get confused if I try to do so in too much detail. Rather, I try to focus on the big message, and work through the figures, which tend to contain the meat.

    ReplyDelete
  10. Stephen Lee

    Class 14, Assignment 14b
    Phenotypic Mapping of Metabolic Profiles Using Self-Organizing Maps of High-Dimensional Mass Spectrometry Data

    (0) Knew:
    I have foundational knowledge of metabolic systems, though I am limited in my knowledge of metabolomics data analysis (MEDI heat maps, etc).
    (1) Learned:
    For one, I learned about the structure of self-administration experiments to develop behavioral models. I learned that data used to generate MEDI heat maps can be used to isolate unique features that distinguish experienced from naïve subjects in exposure studies.
    (2) Pressing Questions:
    What is the potential recourse for those metabolites identified that had no database match (no corresponding mass, ion type)? Are these canonically recognized substances, noise, or could they be new contributions to the metabolome?
    (3) Presentation Topic:
    Application of the described methods to other metabolomics platforms (what would the constituent components be?)
    (4) Thoughts:
    This was an interesting paper. Looking forward to seeing how this discussion proceeds in class.

    ReplyDelete
  11. Article 14b
    Phenotypic mapping of metabolic profiles using self-organizing maps of high-dimensional mass spectrometry data

    0 (Knew): I knew about the existence of SOMs, and I helped present on PCA analysis so I was familiar with that in detail. I also have done research that utilized mass spectrometry so I am somewhat familiar with its use in terms of instrumentation.

    1 (Learned): I learned about the process by which rats are addicted to cocaine, and about the basics of MEDI.

    2 (Questions): I think something that I'm not grasping about MEDI is this phrase: "centroiding and aligning retention time." I think I somewhat grasp it but will look forward to the discussion so that I can gain a solid command of it.

    3 (Presentation): What Stephen Lee said: A different application of MEDI.

    4 (Comments): I'm not sure why, but the wording of this paper was tough for me to get through.

    ReplyDelete
  12. 0.Knew: I knew about SOMs and was familiar with heat maps from previous heat maps.

    1.Learned: I learned more about MEDI and about how it can be used with LC-MS to see that features that are grouped together can show a variety of chemical properties. Like Stephen, I learned about the structure of self-administration experiments, particularly in the case of rats. I got a little more acquainted with PCA. I learned what biological potential is and, in general, another application of metabolomics. I thought it was interesting that

    2.Pressing Questions: I would like to know more about Collision induced dissociation.(p.6565)
    What is Umetrics extended statistics software?(p.6565)
    The description of OPLS-DA confused me. How does it find the relationship between the UPLC-MS data and the cocaine history?(p.6568)

    3.Presentation: brief review of the data analysis in this paper
    Also OPLS-DA.

    4.Thoughts: I thought the descriptions of the methods used to analyze the metabolomics data were confusing for the most part. I think at times I spent too much time trying to understand novel terms and techniques that I got sidetracked from the overall message. However, I did think the paper was interesting. I felt sorry for the rats.

    ReplyDelete
  13. Sylvia Morrow

    Asgn14B_Cody R. Goodwin, Stacy D. Sherrod, Christina C. Marasco, Brian O. Bachmann, Nicole Schramm-Sapyta, John P. Wikswo, and John A. McLean. Phenotypic Mapping of Metabolic Profiles Using Self-Organizing Maps of High-Dimensional Mass Spectrometry Data.

    0. KNEW: A little about IM-MS from the previous paper and a very general understanding of animal testing.

    1. LEARNED: Improved technical understanding of the order and structure of an integrated analysis (in particular the MEDI workflow). A slightly better understanding of how SOMs work, but would like to know more. This paper was most helpful for me in gaining a better understanding of the direct connection between the information learned from MS, PCR, ect. and pharmacological application.s

    2. PRESSING ?:
    --General mass spec questions: I'm familiar with mass spec in the context of separating nuclear isotopes. I assume this is a pretty similar process. If so, do you need to worry about molecules being pulled apart, or are the bonds strong compared to the magnetic forces? In nuclear isotopes you only have to worry about protons and neutrons (so + and neutral particles). I think molecules consist of both + and - ions(?) If this is the case, won't it put extra stress on the molecular bonds and skew the m/z distribution? Also, as far as differential analysis goes, don't you ruin the sample once you analyze it? Are time points from different samples?

    3. PRESENTATION:
    --SOMs

    4. THOUGHTS: A lot of technical language, but it was easy enough to skim most of that without losing the primary narrative.

    ReplyDelete
  14. 0 Knew
    Familiar with SOMs and MS from last reading

    1 Learned
    How to use MEDI to extract metabolomic features of interest; as Madison noted, it was really helpful to see the MEDI applied using real data

    3 Questions
    Like Ben R, what does this mean "The S- plot graphs features based on group specificity or correlation (ordinate) and covariance (abscissa)."

    4 Presentation
    More on how fig 3c was thresholded (Kaskal-Wallis one-way ANOVA)
    More on MS data processing (why/how–deisotoping, centroidization)
    Other applications of MEDI

    5 Thoughts
    Definitely an interesting paper, and less related to the sys bio of it, I am wondering about the technicalities of conducting an experiment such as this (w.r.t. IACUC, ethics, etc.).

    ReplyDelete