The Systems Biology of COVID-19 and the SARS-CoV-2 virus. The class will build a foundation that includes the emergence of complexity, simple biological subsystems, their reductionist and equivalent toy and organ-chip models, and the measurements required to specify model architecture and parameters. Applications to biology, physiology, medicine, chemical and biological defense, pharmacology, drug discovery, and toxicology. UGrad: PHYS 240 01 and BME 290B; Grad: PHYS 326 and BME 395C.
Tuesday, September 20, 2016
SysBio16 Asgn_9A_Class_09_Article_09_2016_09_22
Gene correlation networks: Read Article 09 A.J. Butte and I.S. Kohane,
Mutual information relevance networks: functional genomic clustering
using pairwise entropy measurements. PacSympBioComp, 5: 415-426 (2000).
Post a PCRC on the Blog. Be sure that you dig into the concepts of
entropy and mutual information!
Subscribe to:
Post Comments (Atom)
0. Knew:
ReplyDeleteI knew about relevance networks from what we have discussed in class and read in the papers.
1. Learned:
I learned about how the entropy of an RNA expression pattern is used to determine the mutual information of gene expression. Qualitatively, the mutual information between two expressions is the measurement of how independent they are compared to each other. For example, a low MI means that there is not a lot of new information if we consider the genes together as opposed to separately, whereas a high MI means that we get more information when considering the genes together as opposed to by themselves (and that these genes could be biologically related).
I learned about how these relevance networks are constructed. This paper focuses specifically on networking genes in Saccharomyces cerevisiae. Every gene was compared with every other gene to obtain a MI for each gene. Then, only genes higher than a predetermined threshold MI value were selected to be in the “final” network. The results section discusses how they determined that the distribution of MI’s were significant, as well as how the network changes based on the chosen threshold MI value. By setting the threshold MI to 1.3, 22 relevance networks were constructed. The rest of the paper goes on to discuss the significance of each network as well as the success of this method of creating these relevance networks.
2. Pressing Questions:
The paper mentions on the third page that the gene data used considered a variety of conditions at various time points in each condition. Does this mean that these relevance networks are not conditionally/temporally specific? Did the networks that were produced relate to specific conditions? I know that the networks were grouped together (same gene, similar functions, etc.), but what about the mentioned conditions (diauxic shift, mitotic cell division cycle, etc.) Would it be beneficial to only look at gene expressions for a certain condition at a certain time? Maybe the four graphs on the fifth page would be different?
3. Presentation Topic:
A closer look into the math such as the entropy, MI, etc. would be nice. I know it would probably be more on the statistical side, but I think it would solidify what a connection between two genes really means.
4. Thoughts:
I really enjoyed this paper. I was able to grasp the main concepts and it gives me more of an understanding of how some of these networks can be constructed.
1) Good summary.
Delete2) The first step it to identify correlation between two genes. Hence it is good to exercise the connections thoroughly. This paper is not specifying the direction of the connection, i.e., causality.
3) Yes, **Entropy and Shannon Information**
4) Good. It is a challenge to figure out how to get up to speed with such a wide field.
Number 2 is a good point. Is there a good unidirectional notion of information that allows you to establish causality, or something like it?
DeleteI wouldn't think so, at least for this model of information. If you apply bayes theorem to the second terms of equation two, you can restate the mutual information in the other direction (What does B tell you about A)?
DeleteChinowsky_TheorSysBio_PCRC_0922016
ReplyDelete0. Knew
I knew about relevance networks from previous discussions, and I know the basics of RNA expression.
1. Learned
This paper looks that RNA expression as a way to identify functional gene clusters, basically grouping genes together by function. I learned that related genes have a smaller Euclidean distance, which allows for the creation of self-organizing maps. However, the authors use a methodology that compares all genes against each other using a metric known as mutual information. This metric can be calculated using the entropy of the gene expression patterns. A threshold (TMI) was chosen, so that the resultant relevance network consisted of gene with strong connections to each other.
2. Pressing
In section 3.3, the author says “a gene-gene association with a high mutual information means the expression of one RNA is predictable given the other”. However, given that we know have a much deeper understanding of RNA expression after the ENCODE project, is this view still valid?
How is TMI chosen?
3. Presentation
How the entropy of gene expression can be calculated. How our understanding of genes and gene expression has changed after the Human Genome Project and ENCODE.
4. Thoughts
This is an older paper, but the techniques described within it seem very sophisticated, and like they would be able to at least identify some genes that might need further research, as new links can emerge from this process. This paper about presented an easy to understand way that networks can be constructed.
1) Goood.
Delete2) My hunch is that encode provides more validated data, but the connectivity is fundamental and represents network topology.
3) **ENCODE Project** Volunteer?
4) Yes, a good paper.
Ben Terrones
ReplyDeleteSysBio16 Asgn_9A_Class_09_Article_09_2016_09_22
0. KNEW:
I knew about relevance networks from in class discussions.
1. LEARNED:
I learned about the different methods to cluster genes and show connections between them including the topic of the paper, mutual information relevance networks. I learned what entropy in an RNA expression pattern is and how mutual information (a higher MI between two genes means one is non-randomly associated with the other) can be measured from gene entropies. Using the entropies and mutual information between genes they constructed Relevance Networks to show gene connections. By using a threshold to remove some links they were able to find biologically relevant clusters of genes. They also discussed how changing the TMI would change the resulting clusters. They went on to discuss the strengths of these networks and how they are moving toward temporality and using the technique to identify possible therapy targets in disease physiology.
2. PRESSING ?:
How exactly would a correlation coefficient be calculated and how does that differ from the calculation of entropies and mutual information?
3. PRESENTATION:
Other techniques for creating gene clusters in more detail, specifically self-organizing maps.
4. THOUGHTS:
This was a very interesting paper that was for the most part easy to understand.
1) The concept of thresholds is important!
Delete2) ** Pearson vs Spearman ** Volunteer? May not be critical
3) Will get to this.
4) Good
Kelly McGee
ReplyDeleteArticle 9a
MUTUAL INFORMATION RELEVANCE NETWORKS:
FUNCTIONAL GENOMIC CLUSTERING USING
PAIRWISE ENTROPY MEASUREMENTS
0: Knew: I knew that single gene knockout studies were not extremely helpful in developing an understanding of gene function, as it is a highly multidimensional problem. I was familiar with many of the statistics terms the authors threw around, although I can't say I fully understood the mathematical basis for everything they did.
1: Learned: I learned about Intervention/Calculation, SOM, and metric gene function studies. I learned about the probability-based method the authors used to create nodes of function-related genes, and how it is a useful, manipulable (through TMI)method of discovering associations of different genes through functionality.
2: Questions: Where specifically are they deriving Equation 1 from? Obviously it did what they wanted it too, and some of the correlations they found certainly made sense, but I would like to know the mathematical basis for the equation and the mathematical logic behind its application here.
TMI is obviously variable and arbitrary. The authors did an excellent job of explaining what happens when it is lowered or raised. My question is, what relative values of TMI would be useful for various studies? I.E. what kinds of research questions would best be answered through different values of TMI?
3: Presentation: Any other study that has used this method, and how they set the TMI.
4: Comments: This was exactly the kind of paper I was looking for after Tuesday's "Algorithmic Complexity" debate.
0) High dimensionality is the name of the game
Delete1) OK
2) This should be covered in class.
3) We need to see who referenced this paper and expanded upon it.
4) Good!
0 KNEW
ReplyDeleteI knew about relevance networks, but mutual information was a new concept for me.
1 LEARNED
The concept of entropy as it applies to gene expression was new and interesting to me. The framing of multual information in terms of conditional probabilities was also conceptually intuitive.
The permutation analysis was also very cool control (in fig 1) found this webpage which explains the basic approach.
https://biowize.wordpress.com/2013/12/05/permutation-testing-for-differential-expression-analysis/
One key question in this analysis and any other is "where do you draw the cut-off for connection?" The permutation analysis provides a very sensible way to determine the correct TMI.
I also liked their discussion about they "types" of links which emerged. It makes a lot of sense that identical copies of the same gene are "linked." I would think any good model should reveal the same links.
2 QUESTIONS
Is information entropy "real" in biological systems or just a more general approach to measuring covariation than correlation?
3 PRESENTATION
I'd like to talk about "controls" in systems biology, what are you comparing your model to? In my old lab, one approach was to create a random network with the same degree distribution and to compare the network structure of the random network to the real world one. This paper takes a similar approach to determining a meaningful TMI cut-off. Are similar controls always necessary and how do people design them?
4 THOUGHTS
(Probably biased since I choose this paper) I really liked this paper. One thing to note is that its been cited 846 times, lots of systems-biology work is built upon this work, which appears so simple.
1) I also thought the permutation comparison was a really cool way to find a cutoff! It makes there TMI value much less abstract and it really helped me to understand how the pair-wise measurements actually showed relationships.
DeleteI liked the approach as well.. is this a common practice for investigating expression correlations?
Delete0) We need to learn more about mutual information
Delete1) Excellent link! Thanks. This is a powerful way to distinguish real from random correlations.
2) We will discuss information entropy in class. It is real.
3) Good question - let's see what comes out of the discussion.
4) Good choice!
Natalie Hawken
ReplyDeleteAsgn 9A, Article 09: Mutual relevance networks: functional genomic clustering using pairwise entropy measurements
0. KNEW
I knew that microarrays were used to obtain expression levels of RNA and that this information could be linked to the expression levels of the corresponding genes. I knew what Euclidean distance was, but I had only ever seen it used in a max of three dimensions. I knew that genes could be linked/clustered together for having similar function or for being in the same pathway.
1. LEARNED
I learned that expression levels could be clustered in multiple ways: simple correlation, Euclidean distance, and pairwise comparisons. I learned that entropy refers to the random distribution of expression levels (higher entropy = more randomly distributed) and that mutual information is the measure of additional info known about one gene expression pattern when given another (how well they are associated). I learned that changing the threshold MI changes how many cluster form, how well connected the nodes are, and how significant each part of the network is.
2. MOST PRESSING QUESTIONS
Based on the idea that the expression levels are directly associated with each other, how can we account for genes that inhibit each other (when high expression of one gene leads to low expression of another)?
(Sect 2.3) In 22 networks, 199 genes were connected. Were there any significant connections studied in the canonical network for S. cerevisiae that this model missed?
3. PRESENTATION TOPICS
(Sect 3.2) What is a Parzan density function?
Has the group tested out their goal of including temporal association (sect 3.3)? If so, how successful was it?
4. THOUGHTS
It was interesting to see how these network maps can be created from a large set of data points for hundreds of genes under lots of different environmental conditions. It makes a lot more sense how the RTA project can get such a large and significant network of genes. I appreciated this paper for not having any complex mathematics that would totally go over my head.
I was going to look it up, but I'll just ask here.. what exactly is Euclidean distance? If too complicated to explain, I can Wikipedia..
DeleteSame question about inhibition.. These networks don't necessarily show what the relation is, just that they're related, right? similar to what we've seen with other approaches
So if you have two points (point A at (x1, y1) and point B at (x2,y2) in two dimensions (like a simple graph with X and Y axes), you use the formula SQRT[ (x1-x1)^2 + (y1-y2)^2], to give you the straight-line difference between the two points. But, you can expand this formula for as many dimensions as you need to. You just add to the inside of the sqaure root the square of the difference for each dimension. (For three dimensions, the formula would be SQRT[ (x1-x1)^2 + (y1-y2)^2 + (z1-z2)^2].
DeleteEuclidean distance gives us a numebr that we can use to compare how similar/close the points are. For networks, the smaller Euclidean distance will mean that the genes are more closely associated.
With the relation port, I don't think the network will show inhibition because it seems that the association is dependent on direct proportionality between the two genes (not inverse proportionality too).
0) Natalie explains Euclidean distance. High dimensions are important - see me if you are interested in analytical geometry in high dimensions. **Gunnar Carlson at Stanford**
Delete1-2) Correlation in this context includes anticorrelation. The issues are linear vs monotonic. This paper addresses value of MI for systems with outliers. Inhibition requires causality, for which the network must be directed. That is harder than simple connectivity.
2) Connection to KEGG-type network will be important. Casual mention in this paper. Worth seeing where this group has gone in 16 years.
3) Temporal association is VERY IMPORTANT!
4) Glad you liked the paper.
0.Knew: This was an article filled with more things I learned than things I knew. I did know, however, that entropy was a measure of the randomness or disorder in a system.
ReplyDelete1.Learned: I learned about the three categories of current methodologies in functional genomics that use RNA expression data for creating clusters. I learned about the equation used to calculate the entropy of an RNA expression pattern, and the paper even described what the variables of the equation referred to. When I first read the term “mutual information,” I thought it would be the information shared between two genes, but it is actually the additional information known about one expression pattern when given another pattern. Still, it makes sense that a biological relationship is more likely between two genes with a higher mutual relationship. I learned that relevance networks clusters of genes that are more strongly connected to each other than the threshold mutual information. I learned how the TMI 1.3 was chosen for this relevance network. Higher TMIs would include only the strongest hypothetical associations between genes, and lower TMIs present too many new hypotheses.
2.Pressing Questions: The paper mentioned that the methodology involving the construction of phylogenic-like trees is used to find five waves of expression during embryonic neural development. What would these fie waves be? (p.2)
What are cell state transition rules? (p.2)
Genes were measured under a variety of conditions, including diauxic shift and reducing shocks. What are these two conditions? (p.3)
I am confused a little about how networks can link identical genes. (p.7)
3.Presentation: The description of self-organizing maps (SOMs) makes sense, but it is a little difficult for me to visualize it. Maybe we could see some different examples of SOMs.
4.Thoughts: I like how the article explains the equations used, and I like how it was just lengthy enough to explain things like genetic clusters and relevance networks. It was not too difficult to understand.
0-1) OK
Delete2) Good questions - the paper presenters should address them!
3) We will look into SOMs some more.
4) Good
0. Knew:
ReplyDeleteI wasn't too familiar with the topic before reading the paper, other than what we've discussed previously. I knew that investigating functions of genes was of interest in the field and of the few approaches we've discussed in class.
1. Learned:
This approach was new to me and I learned a lot of how one could use entropy to look relationships in gene function. I didn't know that running perturbations on the genes and re-checking was a way to look for significance. Difference between using mutual information vs correlation coefficients for investigating gene expression. Scale of number of relevance networks can be created using this technique (I thought it would be much larger than 22 networks.
2. Pressing ?:
What would these networks look like when "temporality" is considered? Seems like this will get complicated and may be difficult to actually use the information.
3. Presentation:
How the RNA expression was measured for this particular study (reading/discussing citation 11) and how those measures were input to this system.
4. Thoughts:
This paper was concise and to the point which made it very easy to read and follow. I like the idea but definitely think the future directions will be useful.
0. KNEW
ReplyDeleteI am familiar with methods to obtain RNA expression data, such as microarrays, RNAseq, and SAGE. I'm also familiar with the clustering concept, but not necessarily how they're made.
1. LEARNED
There is a method called Relevance Networks that 'takes large data sets and ascertains facts by performing pairwise correlation coefficients.' This method can take large data sets of RNA expression measured under various conditions and generate networks of hypotheses of gene-gene interactions. They use terms called 'entropy' and 'mutual information' where entropy is an RNA expression pattern is a measure of the info content in that pattern and mutual info is calculated from binary measurements of gene expression. Higher entropy corresponds to more randomly distributed gene expression levels and higher mutual information corresponds to a higher likelihood that between two genes, one gene is non-randomly associated with the other.
2. QUESTIONS
Why does permuting RNA expression form a basis of comparison for the pair-wise mutual information of RNA expression?
Has this TMI-based Relevant Networks model been improved on since it was published in 2000?
3. PRESENTATION
An overview on where equation 1 came from and relevance of skew distributions over normal distributions would be nice to talk about.
4. THOUGHTS
It's convenient that mutual information does not require an a priori choosing of any particular model, for then it lends itself merit as a valuable method as no assumptions need to be made.
Sylvia Morrow
ReplyDeleteAsgn9A_A.J. Butte and I.S. Kohane, Mutual information relevance networks: functional genomic clustering using pairwise entropy measurements.
0. KNEW: Familiar with pairwise correlations from UHI collisions. I don't think we called them Relevance Networks before, but the node-based mapping was familiar from previous papers.
1. LEARNED: About entropy of gene expression patterns and mutual information. I think it's great when you can take well studied tools (like statistical mechanics and entropy) and apply them in places they weren't originally designed for. The logic behind choosing a value to consider "significant" change. The pairwise correlations were used to show that you could drop the TMI from 2 to 1.2. By counting the number of relevance networks as TMI changed, it started to make more sense to me how one might go about isolating a module.
2. PRESSING ?:
--I'm having a hard time finding an intuitive way to understand why connectivity follows the trend seen. The drop around TMI=1.2 makes sense but I wouldn't expect it to go so low, and the peaks at TMI=0 and 2 make sense but I wouldn't expect them to go so high.
--Would there be value in looking at a differential change in # of Relevance Networks to get clues for how things are connected?
3. PRESENTATION:
--Gene expression 101: I can see how MI could be a thing but don't have a concrete understanding of what form this comes in.
4. THOUGHTS: Great papers. Had some nice answers to things I have been wondering about.
This comment has been removed by the author.
ReplyDeleteStephen Lee
ReplyDeleteClass 09, Assignment 9a
Mutual information relevance networks: Functional genomic clustering using pairwise entropy measurements
(0) Knew:
I know a bit about methodologies in functional genomics (DNA/RNA sequencing, microarrays, etc.) Using this data to do gene “clustering” was a concept that was introduced to me in the RTA presentation. Correlation between expression vectors (as demonstrated by Eisen et al) is an unfamiliar method.
(1) Learned:
I learned about the mathematics of expression patterns and constructing relevance networks. A lot of these strategies were already familiar to me from the RTA presentation, but were nonetheless interesting because of differences in network construction and visualization (such as distinguishing stronger/weaker associations of genes).
(2) Pressing Questions:
How confident can we be that genes above the threshold for mutual information actually maintain a functional biological relationship? Could they also represent physical adjacency, regulation by similar transcription factors, etc?
(3) Presentation Topic:
Validity of TMI as a tool for evaluating gene clusters
(4) Thoughts:
This was an interesting paper. Looking forward to seeing how this discussion proceeds in class.
(2)--Exactly. But I think this question is tied to what we're looking for in the study: what kind of correlation relationships are we looking for? I think that question will help us set, or at least understand, the TMI setpoint.
DeleteAgreed. There should be some rationale behind a TMI threshold, and that could potentially be provided by weighing associations already present in the literature.
Delete0 Knew
ReplyDeleteI was familiar with node-based mapping and relevance networks.
1 Learned
The merits of mutual information Relevance Networks in comparison to using correlation coefficients: more general and less likely to mischaracterize functional relationships, uses beyond genomic clustering. Entropy as a concept in gene expression.
2 Questions
Since TMI is assigned somewhat arbitrarily, how do you compare TMI across studies?
3 Presentation
Temporality in Relevance Networks
4 Thoughts
I enjoyed the framing of the paper (mutual information using conditional probabilities), it was intuitive and well-paced. Reading section Section 2.3 (Relevance Networks Seen in Saccharomyces cerevisiae) made me wish I was more familiar with the particular genes referenced (& in general), seemed like a very interesting discussion.
Class 9. Article 9A: A.J. Butte and I.S. Kohane, Mutual information relevance networks: functional genomic clustering using pairwise entropy measurements.
ReplyDelete0. KNEW:
I am familiar with thinking about entropy as a way to quantify things like the information and connectivity within some system. For example, such an approach has recently become fashionable in the real-life-divorced part of theoretical physics to facilitate the study of black hole information loss, and the relationship between quantum mechanics and gravity. It seemed like a reasonable leap to apply it to other sorts of data analyses.
As we have been talking about (and as has been emphasized extensively in the genetics module of my IGP course), it is not enough to sequence a genome. You need to know what the hell each part of it does!
1. LEARNED:
I learned some details about how one might implement a particular clustering algorithm. To be honest, this is simple enough that we could have come up with it ourselves after some thought. The key thing is that you just need a way to measure the 'distance' between two genes, and in this article they decided to use the notion of 'mutual information' to do that.
2. PRESSING QUESTIONS:
The method seems to be straightforward, and it seems like it works reasonably well. I have seen similar things before. One natural question is: does there exist a notable case where this or similar methods have failed spectacularly? If one does exist, what does that mean?
3. PRESENTATION:
Someone should start by saying that you'd like to group related genes/things/whatever together, and that one way to do that is to introduce some notion of distance between things. Then they should motivate why mutual information is a good notion. Then maybe show an example of that the approach and the kinds of graphs it generates. And perhaps an example of when it has discovered something nontrivial.
4. THOUGHTS:
I am glad there were equations in there papers. That means I already understand everything.