Mass Spectrometry: Read Article 16 Jody C. May and John A. McLean.
Advanced Multidimensional Separations in Mass Spectrometry: Navigating
the Big Data Deluge. Annual Rev.Anal.Chem. 9 (1):387-409, 2016.
This is a good review of the cutting-edge of mass spectrometry,
particularly in the context of systems biology. Read it carefully, and
post a PCRC.
Chinowsky_TheorSysBio_PCRC_101116
ReplyDelete0. Knew
As with the other paper, we had previously discussed self-organizing maps in the context of other papers, and I’ve read previous papers on mass spec for other classes. Additionally, our discussions in class about PCA were very helpful in the context of this paper. The information about use of SOM in MS data analysis from the other paper was also helpful
1. Learned
This paper begins with a discussion of big data, and the challenges it represents to the scientific community as a whole, given the “Big Numbers” once must deal with in these data sets. Briefly, the authors discuss both data crowdsourcing and cloud computing, which have been used to deal with other type of data sets in the past. The paper then moves on to talk about different types of multidimensional data methods that are based in mass spec, and also discusses how mass spec is well suited for big data, especially when combined with orthogonal information, which further allows us to reduce complex problems into manageable data sets. The last piece of this paper is a discussion on strategies for the visualization of big data; one of these is SOM, which we have already seen an example of, but they also mention the use of cloud plots, which feature an interactive graphic interface and allow a visualization of a “metabolic bubble”
2. Pressing
As we have read, SOM is the most common way to deal with big data sets. However, I am also curious about cloud plots, and if there are situations where a cloud plot would be better suited than SOM?
3. Presentation
This is a relatively recent paper, so perhaps a quick overview/history of big data, which seems to have appeared very rapidly and became very important just as rapidly.
4. Thoughts
This was a pleasant paper that accurately illustrated the challenges of working with big data, but also made it clear that big data presents nearly endless opportunities for new innovations—just at this point, we are generating it far faster than we can interpret it.
0) Good
Delete1) Well stated
2) Ask Stacy
3) We should discuss this in the context of RTA/Bendamustine
4) Yes - excellent opportunities. A big challenge.
0. Knew:
ReplyDeleteKnew a little bit about MS and the data that can be drawn from MS analysis. Familiar with the "big data paradigm" and SOMs.
1. Learned:
I didn't know there were so many catalogs for possible molecules and databases of known data. Learned about all of the different MS techniques (would like to learn more about these) and how the data is directly translated to SOMs.
The different mass defect scales used in MS were new to me and learned a little bit about CCS and the data produced from CCS. Cloud plots were unfamiliar to me before this paper.
2. Pressing ?:
Didn't quite understand Figure 5 and the mass difference reading from MS in general. I think this question could be answered by a brief overview of MS basics.
3. Presentation:
Basics of MS and brief run through on differences in hybrid MS techniques
4. Thoughts:
Like Colbie said, easy to read and did a nice job explaining the challenges of the "big data paradigm."
1) Lots of catalogs - their annotation is VIP! Cloud plots are neat.
Delete2) Be sure to ask Stacy!
3) We will get an overview by students in the class.
4) Yes!
0. KNEW:
ReplyDeleteI was familiar with MS, but did some more research beforehand to better grasp the technique. I was also familiar with the big data paradigm and how it relates to systems biology. I knew about the different omics sciences from class discussion, but gained a better idea of the scope of information within genomics, proteomics, and metabolomics.
1. LEARNED:
Even though I knew about the big data paradigm, I learned about the different contexts where this is apparent based on the histograms in figure 1. Even though I do not think all of these are relevant to the paradigm (chemical elements discovered), it does a good job of illustrating the idea. I also learned how MS technology is well-suited to face the challenges of big data. I had no idea that MSI and MS experiments could generate 10 billion peaks at 1 million peaks per second, or anywhere near that magnitude. I also learned how cloud maps and SOMs can be useful in visualizing large datasets.
2. PRESSING QUESTIONS:
What are other fields that are dealing with the big data paradigm other than systems biology, and how are they approaching the vast amount of information? Will the limiting factor in understanding the information be the computational methods or the human involvement in understanding the information that is obtained? Especially when looking at multidimensional spaces, there are many ideas that are harder for humans to comprehend than computer models.
3. PRESENTATION:
MS basics with details on multidimensional and hybrid techniques. Figures 5&6 could be clarified from class discussion.
4. THOUGHTS:
I liked how the paper finished with an SOM of a liver-on-a-chip exposure to APAP because it was a nice way to tie back concepts that we previously learned in class. The paper did a good job putting context around the enormity of data that is being handled.
0) Good
Delete1) Good illustrations!
2) As the article said, particle physics, astronomy, genetics
3) Will have a presentation
4) Good
Class 14. Article 16 Jody C. May and John A. McLean. Advanced Multidimensional Separations in Mass Spectrometry: Navigating the Big Data Deluge. Annual Rev.Anal.Chem. 9 (1):387-409, 2016.
ReplyDelete0. KNEW:
I was aware of mass spectrometry as a useful tool which identifies the chemical constituents of something using (probably unique) charge-to-mass ratios. I was also aware, based on some of Dr. Wikswo's comments, that the method is an important one in understanding mechanisms of action, partly because it is fast, can screen a lot of things, and works. I am also broadly aware from the class of many of the challenges associated with a flood of data (like data visualization and identifying patterns more generally) which must be interpreted and utilized somehow. Finally, we have talked about self-organizing maps and principal component analysis in class, and I think I have a good handle on them.
1. LEARNED:
The big thing I got from this paper was a change in perspective; I never thought about there being a large database of compounds (PubChem), and how that might help reduce the problem of identifying compounds to making measurements and comparing them with that database. To me, this seems to imply that you don't need a lot of actual knowledge of chemistry to do state of the art measurements: you just need a computer and a sufficiently up to date database.
I also learned of the efforts which apparently exist in the mass spectrometry community to make dealing with data more manageable. As I mentioned in the previous section, while I did not know about this, it seems plausible.
I also was reminded of the kinds of measurements which may complement mass spectrometry data; for example, ion mobility experiments (separating things based on size) and mass defect analysis (perturbations of different chemical compositions will be different in general, even if the two original masses were similar).
2. QUESTIONS:
My big question has to do with what I pointed out in 1: namely, must one know chemistry to do chemistry? It seems like with advanced enough databases and measurement tools, chemical knowledge becomes unimportant (at least in these particular experiments). Can mass spectrometry become completely automated with a good enough algorithm, so that human involvement becomes completely unnecessary? One can easily imagine an algorithm which references the PubChem database to identify compounds, and uses data analysis tricks like self-organizing maps to identify trends in data.
3. PRESENTATION:
A good presentation would explain mass spectrometry in detail, compare it with other methods mentioned in the paper, and explain why it might be a good choice for attacking systems biology-related problems. There should also be some mention of the challenges associated with a deluge of mass spectrometry data, which is a main focus of the paper.
It would be helpful to see an explicit example of data being analyzed this way.
4. THOUGHTS:
I get the impression that, once again, one does not need to know much chemistry to be able to do research in a mass spectrometry lab. I wonder how much of this impression is true, or if I am just being arrogant/naive.
0-1) Good things to know and learn
Delete2) The key question is the extent to which the knowledge of chemistry/biology is needed to interpret the results. I believe that both are critical.
3) Yes
4) The issue is how far can numerology alone take you!!
0. KNEW
ReplyDeleteI have a general understanding of mass spectrometry as a means to determine compound composition, but not much else in regards to how it's applied in systems biology. I am aware of the existence of the omnics from class.
1. LEARNED
Proteomics: Detect/measure all proteins found in an organism.
Metabolomics: Pertaining to small-molecule metabolites.
Genomics: Pertaining to genes.
MS is suited to big stat as the throughput and info density are high. While MS directly provides the exact mass of the analyst, mass defect permits locating related chemical species in a complex MS spectrum. It helps to identify exogenous metabolites that have chemical comps related to the drug species of interest.
Methods to visualize big datasets include cloud plots, SOMs
2. QUESTIONS
How important is it to obtain direct, quantitative information regarding analyst similarity? It seems to be a big drawback to principal component analysis. (p. 401)
3. PRESENTATION
Mass spectrometry basics.
4. THOUGHTS
A nicely written paper. It elucidates why MS is helpful, even though some of the jargon went over my head. I'm getting a better idea as to how SOMs and PCA tie in with MS.
0-1) OK
Delete2) This is a serious problem with PCA - it is good for decisions, but not necessarily deep insights. WE NEED TO DISCUSS EFFECTIVE MODELS.
3) Yes
4) Agree
Ben Terrones
ReplyDeleteSysBio16 Asgn_14A_Class_14_Article_16_2016_10_11
0. KNEW:
I knew that one of the biggest challenges in systems biology is the amount of data that is being dealt with. I knew what self-organizing maps were from previous papers and class discussions.
1. LEARNED:
I first learned about the three V’s used to identify a big data challenge: large volume of data, high velocity, and a variety of different subsets. I knew what crowdsourcing was but I hadn’t heard about it being used for scientific purposes. The authors then go on to talk about how MS is a good fit for big data because of its high throughput capability, and they show this with the peak capacities of various MS-based techniques. They then discussed how using multidimensional MS, it is possible to assign a structural identity to an unknown analyte. Mass defect analysis is used to identify small mass differences by rescaling the mass axis using different molecules. At the end of the paper they discuss two methods for visualizing big data: cloud plots and self-organizing maps.
2. PRESSING ?:
What is the gas-phase collision cross section?
3. PRESENTATION:
Presentation on cloud plots would be interesting.
4. THOUGHTS:
I enjoyed this paper; it was for the most part easy to understand and very interesting.
0) There is also a big issue regarding what experiments to conduct - big experiments vs big data.
Delete1) OK
2) Google it!!! Or ask in the MS class.
3) Yes
4) Good
0 Knew
ReplyDeleteGeneral awareness of MS–that it exists, but otherwise very little. Familiar with SOMs and PCA in sys bio and use of SOM differential analysis to extract temporal/analyte specific info from class. Also familiar with various data visualization techniques/design concepts and theories (BS BME & Design)
1 Learned
That there are efforts to crowdsource MS data (was aware of FoldIt and Penguin Watch). Learned about various multidimensional methods based on MS and orthogonal separation dimensions (i.e. mass defect analysis, IM, mobility mass correlations). Meaning of "tractable."
2 Questions
How do molecules overcome valence and chemical stability rules (w.r.t to 200,000 validated structures existing in a mass range for which only 5,000 are expected)? pg 395
How/will the shift in database distribution affect field development in some endogenous way?
Is avergine scale even relevant anymore? Based on 1995 PIR and "axis rescaling is not necessary." (pg 397)
3 Presentation
"Nesting of analytical timescales" w.r.t. MS (pg 396)–Nicole touched on this in RTA presentation
Artificial neural networks
4 Thoughts
150 references–I wonder what they use for a reference manager (Mendeley or something else). Also wondering, similarly to John, when/if proteomic/metabolomic analysis becomes a Big Data problem more so than chemical. I suppose it will require researchers who are adept in multiple phase spaces who can understand the underlying principles of both fields.
0-1) OK
Delete2) ASK!! The issue is possible vs probable vs actual vs useful!
3) We will need to discuss time scales when we revisit Bendamustine.
4) Good questions.
0. KNEW
ReplyDeleteI was familiar with the concept of cloud computing from previous classes as well as independent interest in the subject. I was also aware of crowdsourcing examples, although I was previously unaware of the title crowdsourcing. For example, my high school biology teacher explained the mechanics of protein folding with the online game Foldit.
1. LEARNED
Striking fact that the paper brings up early on includes the capability of high-throughput screening to identify 100,000 compounds per day from canonical libraries as well as the vast genomic data that is projected to be generated in the next decade. The paper discusses three challenging factors to big data—volume, velocity, and variety.
The paper then talks about generating chemical knowledge, noting that over the past fifty years, chemical element discovery has increased relatively linearly; however, chemical substance registration as well as chemical abstract indexing have increased exponentially. Translating chemical information is largely limited on informatics tools, and data accessibility will provide new opportunities to interpret these databases via crowdsourcing and cloud computing. Crowdsourcing is the intended synergistic combination of combined efforts of a large range of people toward solving a common problem, and cloud computing decentralizes data-intensive computational work in order to decrease the need for institutions to invest in instrumentation for data analytics. I also learned about the process and importance of multidimensional mass spectrometry in analyzing the chemical structure of a sample compound.
2. PRESSING ?
What is the range of institutions capable of accessing and interpreting these data sets today, given that cloud computing is an ongoing process to decrease investment of data analytics instruments?
3. PRESENTATION
peak production, peak capacity in IM, LC, MS
4. THOUGHTS
Interesting paper…I liked some deviation to broader topics, such as crowdsourcing
1) Good summary
Delete2) Ask Stacy. Vanderbilt is at the forefront
3) OK - ask. Fig 2 is VIP
4) Good
Assn 14a, Article 16: Advanced Multidimensional Separations in Mass Spectrometry: Navigating the Big Data Deluge
ReplyDelete0. KNOW
Based on the discussions about RTA, I knew that MS analyses resulted in huge amounts of data in pretty short periods of time, and that it can be difficult to parse through this data. Previous papers have also talked about the "leibniz" of possible pathway connections to all of the different post-modification proteins, metabolites, etc, which leads to difficulties in data collection. I knew that mass spec was often done in conjugation to other techniques (chromatography, etc), which increased the amount of data provided. We also had previously discussed SOMs in class.
1. LEARNED
I found Fig 1 to be super interesting. I knew that most computing resources were exponentially increasing in power, but I had no idea that the number of chemical substances entered was also increasing exponentially. I had heard briefly about cloud computing and crowdsourcing, but I had no idea that it could be applied to data analysis for science. I thought the workflow in Fig 3 was really cool, and it helped me to understand how this seemingly simple analysis technique could so specifically identify compounds. I had never heard of the mass defect measurement in MS, and I'm still a bit confused by it. It seemed interesting that people were able to glean even more info from this defect. Fig 6 really showed how these multidimensional approaches improve classification and separation of data sets.
2. MOST PRESSING QUESTIONS
What is the type of algorithm used to develop an SOM? How does the initial separation of the SOM compare to the SOMs made with other data sets?
Where/how is this huge amount of data stored?
3. PRESENTATION TOPICS
Overview of the different types of multidimensional MS (as seen in Fig 2, pg 393). What data does each of these techniques provide? Which techniques are used the most?
How does the Kendrick mass defect provide additional data? What exactly is this "binding energy of nucleons"? (pg 397)
4. THOUGHTS
This was a good overview and introduction to multidimensional MS. I found the different ways of representing the multidimensional data interesting (especially the cloud plots). I'm still working to understand SOMs. Also, because I was bored and curious, I downloaded and starting playing the Foldit game described in the crowdsourcing section, and it's pretty fun (and a good way to procrastinate).
0) OK
Delete1) Good- ask about mass defect
2) We may need an SOM project
3) Ask about mass defect.
4) Glad you are trying Foldit - report on what you found!
0. Knew:
ReplyDeleteAll I knew about MS is that it can determine the intensity of mass/charge ratio of molecules in a certain chemical. I knew that many areas of biology experience/have to deal with big data. I knew about crowdsourcing and the game FoldIt.
1. Learned:
I learned the three V’s of big data challenges: volume, velocity, and variety.
I learned that there are many areas where our knowledge/innovation in dealing with knowledge has been increasing almost exponentially (figure 1). Cloud computing and crowd sourcing are two ways to deal with processing the large amount of data we have.
The main goal in dealing with large data is “to reduce a complex problem into manageable subsets of data.” In MS this is done for molecular identification by using mass measurements to reduce the data set through some initial characterization. This is illustrated in figure 3.
I also learned about cloud plots. I always thought about different ways one could show/visualize multidimensional data. Cloud plots are a really cool way to do that. I also gained more understanding of self-organizing maps since they were briefly discussed in class.
2. Pressing Questions:
Is there a limit on how large our knowledge/processing can become? In other words, won’t there be some point where the graphs in figure 1 stop growing exponentially? I know the paper mentions that the information in 1a is now increasing at a relatively constant rate, but what about the rest? In particular for the number of transistors in a microprocessor or data disk storage capacity, is there a limit to how much data we will be able to hold or process in the future?
3. Presentation Topic:
More in depth presentation on SOMs. I feel like I am almost there in full understanding them, I just need a bit more information.
4. Thoughts:
I originally figured MS was a simple thing that just gives charge to mass ratios of molecules. I enjoyed learning new things about MS and its applications.
0) OK
Delete1) I personally need to get more comfortable with cloud plots.
2) Machine learning is the key. Watch out for the kurzweil singularity!
3) We need an SOM group!
4) Good!
0 Knew
ReplyDeleteI was generally familiar with MS from my days as a chemistry major, and I remember something about exact mass corresponding to the molecular formula. I have also encountered SOMs in class and in Bachmann lab work.
1 Learned
The figures in this paper were great; they did an excellent job of framing the role of big data in MS. Figure 6 was also very interesting, they were able to seperate out the different classes of lipids with the same mass.
2 Pressing Questions
Why do we have to regard increases in abundance serpeate from decreases? Would our maps not be more robust if they were combined?
3 Presentation Topic
More on SOMs.
4 Thoughts
Good review of MS. Two challenges in big data are identifying trends and visualizing multidemensional data. Are these two challenges the same? Can we arrive at the same set of unique features without the visualizations or will we always be dependent on a human observer to define the "important" part of the map.
0) OK
Delete1) Yes
2) Hard to see - a matter of color bar. Has been the subject of some discussion.
3) Yes
4) Let's see. People need help!
0.Knew: I knew from previous readings about self-organizing maps, and I knew that generally there is an exponential increase in chemicals discovered and indexed over the years. I was somewhat familiar with crowdsourcing, and I had heard of Moore’s Law before.
ReplyDelete1.Learned: I learned about the big data paradigm, and I learned that the volume of genomics data is expected to surpass the data of such popular social media sites as YouTube and Twitter by 2025. I learned about cloud computing. I learned that the Chemical Abstracts Registry has been around for about 51 years. One of the more interesting things I came across in this paper was about the Foldit video game, a crowdsourcing method to address protein folding. I learned what an analyte is, and I learned about mass defect analysis (I like how the Kendrick mass defect plot was explained). I learned about cloud plots to visualize big data.
2. Pressing Questions: How many combinatorial small-molecule libraries are there? (p.388)
What exactly is SciFinder? (p.389)
In the section about the crowdsourcing where citizen scientists collected soil samples and screened them through liquid chromatography-mass spectrometry, it is stated that a novel fungal metabolite, maximiscin, was discovered. Maximiscin has exhibited antitumor activity in a mouse model. What is the current state of maximiscin testing now? Any closer to using this to treat tumors? (p.391)
Still a little bit unclear on what is meant by “orthogonal” pieces of information. (p.394)
3. Presentation: review of figures used, especially fig 2 which I found to be a little challenging to interpret.
4. Thoughts: Overall, I thought this paper was interesting. I really appreciated the definitions in the margins and the fact that a lot of figures were used to help with understanding. Though the topic was complex, the article made sure that for the most part, the concepts were explained so that we regular citizen scientists can understand them.
0-1) OK
Delete2) Google SciFinder and Max... Orthogonal - ASK IN CLASS AFTER BREAK!
3) OK
4) It was good article
Sylvia Morrow
ReplyDeleteAsgn14A_Jody C. May and John A. McLean. Advanced Multidimensional Separations in Mass Spectrometry: Navigating the Big Data Deluge.
0. KNEW: Some of the basic structure of big data in the context of RHI experiments.
1. LEARNED: About what experimental methods are available: IM, MS, LC, combinations of these. The method selection process which considers how much data is produced, how quickly data is collected, and how quickly that data can be analyzed.
2. PRESSING ?:
--In nuclear collisions we set up all these detectors as a shell around the collision region so all types of detection happen basically simultaneously. I am interested in whether how exactly these detection methods are combined and what the transfer process is to move from one type of detection to another. If I understand correctly, this paper talked about mass spec mostly in the context of identifying new molecular structures which is different than what we've been talking about which is measuring quantities of known structures.
3. PRESENTATION:
--Novel methods of storing data to reduce size and increase analysis efficiency?
4. THOUGHTS: I love the idea of crowdsourcing via puzzle games.
Stephen Lee
ReplyDeleteClass 14, Assignment 14a
Advanced Multidimensional Separations in Mass Spectrometry: Navigating the Big Data Deluge
(0) Knew:
I have limited knowledge of MS methods and multidimensional configurations for the data gathering described in the paper. What knowledge I do have comes from the presentation on RTA earlier in the semester.
(1) Learned:
I learned about the exponential nature of growth of multidimensional MS both in terms of volume of data gathered and data gathering rate. I learned about different dimensions that can be obtained from MS experiments (such as mass-to-charge, mass defect, mass, etc) to more specifically describe and identify molecules.
(2) Pressing Questions:
Is our data gathering ability outpacing our data processing ability? How will we best use these data with the computational capacity that we have? It seems to me that one of the central challenges in big data is identifying which data is most worth analyzing.
(3) Presentation Topic:
Summary of information dimensions (as mentioned above)
(4) Thoughts:
This was an interesting paper. Looking forward to seeing how this discussion proceeds in class.
Article 14a
ReplyDeleteAdvanced Multidimensional Separations in Mass Spectrometry: Navigating the Big Data Deluge
0 (Knew): I was quite familiar from the first part of this course about the huge amount of data we have to work through in systems biology, and the basics of SOM and PCA. I was also aware of crowdsourcing games for scientific data analysis. I also have some experience with mass spectrometry
1 (Learned): I learned that mass error and collision cross section provide more parameters for higher dimensional analysis of mass spectrometry data, and exposure to a discussion of how best to analyze said data whether by SOMs, PCA, etc.
2 (Questions): Are there analogs to mass error and collision cross section for other instrumentation techniques such as NMR?
3 (Presentation): Crowdsourcing data processes through video games.
4 (Comments): Fascinating paper.