<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Hang Yuan</title>
    <description>Using wearables to improve human health.
</description>
    <link>https://hangyuan.xyz/</link>
    <atom:link href="https://hangyuan.xyz/feed.xml" rel="self" type="application/rss+xml" />
    <pubDate>Mon, 23 Mar 2026 11:11:57 +0000</pubDate>
    <lastBuildDate>Mon, 23 Mar 2026 11:11:57 +0000</lastBuildDate>
    <generator>Jekyll v3.10.0</generator>
    
      <item>
        <title>Being Organic</title>
        <description>&lt;p&gt;Lately, I’ve been reflecting on the concept of being organic. To me, being organic means not forcing things to happen. It’s a mindset of alignment—understanding nature, respecting its pace, and allowing things to unfold naturally. This way of thinking applies not just to personal life, but also to how we work, grow, and create.&lt;/p&gt;

&lt;p&gt;Being organic shares similarities with being authentic—both encourage us to act and speak in alignment with our core values. But while authenticity is inward-looking, being organic extends outward. It includes how we grow over time and how we interact with our environment. It’s less about identity and more about process—about pace, rhythm, and adaptation.&lt;/p&gt;

&lt;p&gt;Being organic means growing not by chasing metrics, but by evolving in synergy with ourselves and the world around us. Society often sets expectations: go to college at 18, work 9 to 5, get married, hit milestones on a fixed timeline. These expectations can be helpful heuristics, but they don’t fit everyone. Some people do their best work at night. Some aren’t ready for college at 18. Some chart entirely different paths.&lt;/p&gt;

&lt;p&gt;Too often, we live as though our stories have already been written, forgetting that the script might not suit us. What we really need, especially early in life or in our careers, is time and space to explore,  to experiment and to grow.&lt;/p&gt;

&lt;p&gt;There are many benefits to living and working organically: greater peace of mind, sustainable growth, deeper alignment with your environment, and better readiness for emerging opportunities. On this last point, I’m reminded of the book &lt;a href=&quot;https://www.google.com/search?client=safari&amp;amp;rls=en&amp;amp;q=Why+Greatness+Cannot+Be+Planned&amp;amp;ie=UTF-8&amp;amp;oe=UTF-8&quot;&gt;Why Greatness Cannot Be Planned&lt;/a&gt;. Its central thesis is powerful: if we reduce our journey to chasing predefined objectives, we limit our creative search space—and often miss unexpected breakthroughs along the way. True progress, it argues, is guided by curiosity and executed with playful exploration, not rigid metrics.&lt;/p&gt;

&lt;p&gt;So, is there only one way to be organic? Of course not. Think of a garden: some plants grow fast, others slow. Some are green and plain, others colourful and intricate. Yet they all grow in the same soil, responding to the same sun and rain, each in their own way. Humans are like that too. Growing organically might mean being as persistent as grass, or as intricate as a flower.&lt;/p&gt;

&lt;p&gt;I’ve been thinking about this in the context of my own research. Recently, I led a small team to develop a generalist medical AI model. We were working under a tight timeline, as I was preparing for a move to the US. The project was ambitious, and the team was brilliant. But the pressure to deliver quickly led us to make assumptions we didn’t have time to test rigorously. I fell into confirmation bias, interpreting weak signals as signs we were ready to scale.&lt;/p&gt;

&lt;p&gt;That was a mistake. In hindsight, I see that the rush cost us insight. We overlooked promising directions because we were fixated on our original goal. Not getting the fellowship I applied for might have been a blessing in disguise—it gave me the time to revisit our assumptions. Now, with a clearer mind and fewer constraints, I believe we’re on track to something more meaningful.&lt;/p&gt;

&lt;p&gt;Can this organic approach scale? I think so—but not through top-down mandates. Being organic doesn’t work by imposition. It grows from the bottom up—from individuals, teams, and communities choosing to live and build differently. Not every challenge has an organic solution; large-scale infrastructure and science projects still require centralised coordination. But in areas where flexibility, creativity, and emergence matter, being organic can lead to a society that is less polarised, less rigid, and more adaptive.&lt;/p&gt;

&lt;p&gt;Be organic when you can. The results will take care of themselves.&lt;/p&gt;
</description>
        <pubDate>Mon, 19 May 2025 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2025/05/19/being_organic.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2025/05/19/being_organic.html</guid>
        
        
      </item>
    
      <item>
        <title>Thinking as a Young Scientist</title>
        <description>&lt;p&gt;&lt;em&gt;Thinking is perhaps the most important and yet most under-valued activity for young scientists&lt;/em&gt;. When I went through my PhD training, I saw myself and other highly motivated peers being mostly occupied with activities of doing science rather than thinking.&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/thinking/think.jpg&quot; /&gt;
    &lt;figcaption&gt;
 Figure 1: A simplified view of the scientific process 
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;

&lt;p&gt;PhD students spend most of their time doing science for good and bad reasons. As trainees, we should learn essential skills that will enable us to do science. Nonetheless, PhD students are often stuck in a cycle of grinding out papers without time for deep thinking. As PhD students, we start a project by reading some papers, forming a new hypothesis and conducting an experiment to test our hypothesis. When an idea works, we write up the results and publish a paper (Figure 1). When an idea fails, we go back to any of the previous steps depending on how much belief we have in the idea itself. The less belief we have, the further up we move in the discovery progress. In a PhD program, we repeat the same loop until we have about 2-3 papers over 3+ years in Europe and 5+ years in the US. By the end of this journey, we bundle up all the results and massage the text to make up a more coherent story for the thesis.&lt;/p&gt;

&lt;p&gt;Thinking as an activity is not built into our workflow because whenever we think. Even though we do spend time thinking when choosing our PhD project and writing our thesis, we will be considered as unproductive in today’s competitive scientific envrionment.&lt;/p&gt;

&lt;p&gt;Thinking in sciences can happen at different layers of abstraction. The day-to-day scientific investigation, domain-specific thinking, and thinking in the context of the history of science:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Day-to-day: thinking related to how to make experiments happen, interpreate the results and refine the hypotheses on a daily basis&lt;/li&gt;
  &lt;li&gt;Domain-specific: how the work done is relevant in the broader context of a discipline in recent development&lt;/li&gt;
  &lt;li&gt;History of science: how the work contributs to the history of science in one or several disciplines over a longer period (10+ years)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Young scientists spend most of their time thinking about their day-to-day investigations and less time on the relevance of the work in domain-specific content. Very little time is spent on thinking about the history of science. I’d argue that thinking at a higher level of abstraction is hugely valuable.&lt;/p&gt;

&lt;p&gt;Due to the lack of domain-specific thinking, most of my machine learning PhD friends left academia because no one will use the novel methods they propose in the real world. They are right because the datasets used in machine learning research are not reflective of the real world. If their models were developed and tested on easier data, why would we expect them to be equally effective outside academia? If the lack of real-world impact is what they value, early in their PhD, they could have first thought about the properties of machine learning problems that have real-world impact, identified the resources they need to work on those problems, and then chose the problems that they have the highest chance of solving. With sufficient thinking early in the PhD training, many such limitations can be addressed. Our schools excel at teaching the students how to solve problems but much less so in choosing which questions to ask. Learning which questions to ask requires deep thinking and is a much harder skill to teach as it is open-ended.&lt;/p&gt;

&lt;p&gt;Many scientists are frustrated by negative results precisely because they have not thought enough about the history of science. The currency in academia is publications. Because of how academia is structured, the chances of getting published are significantly higher if we find positive results. However, most investigations will not produce positive results. The key is to see success through failure, which wouldn’t be possible without thinking about the history of science.&lt;/p&gt;

&lt;p&gt;A good mental model is to consider the accumulation of scientific processes akin to drawing a map of knowledge (Figure 2). For a novel problem, we start with some known facts and then explore different ways of solving this problem (Figure 2.1). In rare circumstances, the first solution we try can address the novel problem, making it relatively easy to get published (Figure 2.2). However, most of the time, we are running in circles, facing dead ends for several years (Figure 2.3). Sadly, negative results are often not published, putting people’s careers on hold. A key point that goes unnoticed is that the negative results can be equally valuable as scientific knowledge. Science is not just about how we solve every problem out there. Science is also about identifying why certain solutions work whilst others don’t. But it takes someone who understands the history of science, aka how science progresses, to appreciate both the positive and negative results. Most of the time, things don’t go the way that we want, but we just need to put trust into the process by filling out the map of knowledge step by step.&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/thinking/map.jpg&quot; /&gt;
    &lt;figcaption&gt;
 Figure 2: The knowledge map of scientific discovery 
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;

&lt;p&gt;Doing science is hard, but thinking about science is harder because it is a less well-taught skill. I know some people like to think in a solitary fashion. For me, at least, I enjoy discussing ideas with my colleagues and, at times, writing down my own thinking to help articulate my thoughts. Every scientist should deliberately spend time thinking on their own and strive to get better at it; otherwise, we might risk thinking superficially, simply reacting to the information we receive. I fear that will be the death of our scientific creativity.&lt;/p&gt;
</description>
        <pubDate>Sun, 15 Sep 2024 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2024/09/15/thinking-as-a-young-scientist.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2024/09/15/thinking-as-a-young-scientist.html</guid>
        
        
      </item>
    
      <item>
        <title>Representation Learning for Genomic Discovery</title>
        <description>&lt;p&gt;About 2-3 years ago, Daphne Koller gave a thought-provoking talk on how representation learning can be used to accelerate drug discovery using genetics at the Big Data Institute, Oxford. I thought her talk was super cool as I also work on representation learning. But I couldn’t really understand what’s been done due to my lack of knowledge in the field of genetics. Thus, I decided to write on this topic to teach myself about the value of representation learning for genomic discovery.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/rl_genomic.jpg&quot; alt=&quot;&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Genetics, environmental exposure, and lifestyle are three pillars of human health. Due to recent advances in sequencing technologies, we can now sequence the human genome much faster and cheaper. The estimated cost of sequencing the first human genome sequencing is 300 million US dollars over 15 months in the Human Genome Project around the year 2000. Today, one can sequence their own DNA at a cost of a few hundred bucks within a few hours. Therefore, large volumes of human genome data have been collected for health discovery at a scale not possible before, such as the &lt;a href=&quot;https://www.ukbiobank.ac.uk&quot;&gt;UK Biobank&lt;/a&gt; and &lt;a href=&quot;https://ourfuturehealth.org.uk&quot;&gt;Our Future Health&lt;/a&gt; initiatives.&lt;/p&gt;

&lt;p&gt;The human genome roughly consists of 3 billion base pairs of nucleotides and 20K genes across 23 chromosomes. To make sense of the functions of each gene, we often deploy a statistical technique called &lt;em&gt;genome-wide Association analysis (GWAS)&lt;/em&gt; by performing lots of logistic regressions to compare whether there are genetic differences across populations. The goal is to determine, for instance, whether populations that have variants of a gene have different body fat or risk of breast cancer. One of the very nice things about genetics is that because our genetic data largely stay unchanged since birth, we identify causal associations between our DNA and traits of interest. Making causal claims will be challenging in other types of observational studies as we will be subject to potential bias and confounding, which is a whole research topic on its own. If you want to know more about causal inference, the &lt;a href=&quot;https://www.amazon.co.uk/Book-Why-Science-Cause-Effect/dp/0241242630&quot;&gt;Book of Why&lt;/a&gt; by Judea Pearl will be a must-read.&lt;/p&gt;

&lt;h2 id=&quot;why-representation-learning&quot;&gt;Why representation learning?&lt;/h2&gt;

&lt;p&gt;Why does representation learning matter for genomic discovery? It matters because GWAS relies on converting the phenotyping measurement into a single scalar value. While the existing approach works for simple phenotypes like height and weight, it will be much more challenging to do so for high-dimensional cross-sectional data such as brain imaging and CT scans or low-dimensional high-frequency data such as wearable sensing data. I will refer to them as &lt;em&gt;high-content clinical data&lt;/em&gt;. When performing GWAS on high-content clinical data, we have the following limitations:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Huge reduction in dimensionality&lt;/strong&gt;: In terms of raw data volume, modern measurement instruments can be high-dimensional. Using wearables as an example, one week of recording can lead to 10M+ data points, but we will have to condense this data sequence into a simple scalar value, such as weekly step count in a GWAS. Regardless of what magic number we come up with, the high degree of dimension reduction will lead to information loss. Even though we can perform the GWAS on every single dimension of the recorded sequence for 10M+ times, it will be computationally intractable.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reliance on expert-curated labels&lt;/strong&gt;: Typically, the phenotype or trait of interest is defined by experts and often also has to be annotated by an expert. Inevitably, expert-defined labels will be limited in volume. So we won’t have enough power to detect the genetic variations that are less common.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Missing features not discernable by humans&lt;/strong&gt;: The high-content clinical data could have subtle features not discernable by humans. In wearable space again, when we look at traces of an accelerometer, it is difficult to know what activity someone is doing. Still, it will be possible to infer the activity being performed using data-driven approaches.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Representation learning&lt;/em&gt; has been investigated to address the issue above for a better-informed genomic discovery pipeline. Representation learning/feature learning describes a class of data-driven learning methods aiming to compress high-dimensional data into a lower-dimension latent space, sometimes referred to as embeddings. &lt;em&gt;Principle component analysis (PCA)&lt;/em&gt; is a commonly used representation. However, PCA only captures linear relationships within its principle components, which is insufficient for high-dimensional clinical data. Figure 1 explains how representation learning-based phenotyping might differ from expert-curated phenotypes. For the rest of the blog post, we will explore how current representation learning methods can help with genomic discovery.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/rl_genomic_over.jpg&quot; alt=&quot;Figure 1: Representation learning for genomic discovery overview&quot; /&gt;&lt;/p&gt;

&lt;h2 id=&quot;how-to-obtain-the-embeddings&quot;&gt;How to obtain the embeddings?&lt;/h2&gt;
&lt;p&gt;Current works mostly rely on using auto-encoder to minimise reconstruction loss. Alternatively, contrastive approaches have also been explored.&lt;/p&gt;

&lt;h3 id=&quot;reconstruction&quot;&gt;Reconstruction&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;REpresentation learning for Genetic discovery on Low-dimensional Embeddings (REGLE)&lt;/strong&gt;&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/rl_genetics/regle.jpg&quot; /&gt;
    &lt;figcaption&gt;
        Figure 2: REGLE overview. Source: &lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/37163049/&quot;&gt;Yun et al., 2023&lt;/a&gt;.
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;

&lt;p&gt;In 2023, Google published one of my favourite papers on phenotype representation learning using reconstruction. The authors proposed a generic framework called REpresentation learning for Genetic discovery on Low-dimensional Embeddings (&lt;a href=&quot;https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10168505/&quot;&gt;REGLE&lt;/a&gt;). The study consists of three steps, shown in Figure 2:&lt;/p&gt;
&lt;ol&gt;
  &lt;li&gt;They learnt the embedding of high-content clinical data including photoplethysmography (PPG) for cardiovascular functions and spirograms for lung functions using variational autoencoder (VAE) as the backbone using reconstruction as the learning objective.&lt;/li&gt;
  &lt;li&gt;GWAS was then performed on the obtained embeddings.&lt;/li&gt;
  &lt;li&gt;Polygenic risk scores (PRS) were computed on each of the coordinates of the embeddings and then used to construct a disease-specific PRS.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What I like about this paper is that instead of using a normal autoencoder, they used VAE instead whose coordinates will be less coupled with each other. Having less correlated coordinates could enforce learning different aspects of the underlying biology.&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/rl_genetics/model_regle.png&quot; /&gt;
    &lt;figcaption&gt;
        Figure 2: Embedding learning design. Source: &lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/37163049/&quot;&gt;Yun et al., 2023&lt;/a&gt;.
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;

&lt;p&gt;REGLE introduced three sets of embeddings, two for spirograms (SPINCs and EDFs+SPINCs) and one for PPG (PLENCs). For spirogram embedding, EDFs+SPINCs embedding set also had the EDFs information injected by feeding the EDFs to the decoder during reconstruction (Figure 2). The embedding-based approaches were able to discover more loci in general for both PPG and spirograms (Table 3). However, it is interesting to note that when embedding was learnt only on spirograms (SPINCs), fewer known loci (510) were discovered than when using the expert-defined phenotypes (581). Not sure if it is because of the larger sample size of spirograms. Perhaps this motivated the authors to add EDFs into the embedding learning to eventually discover greater known loci (596) alone.&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;figcaption&gt;
        Table 1: Comparison of GWAS significant loci. Source: &lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/37163049/&quot;&gt;Yun et al., 2023&lt;/a&gt;.
    &lt;/figcaption&gt;
    &lt;img src=&quot;/assets/images/rl_genetics/regle_loci.png&quot; /&gt;
    &lt;/figure&gt;
&lt;/div&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/rl_genetics/regle_prs.png&quot; /&gt;
    &lt;figcaption&gt;
        Figure 3: Polygenic risk score comparison in the UK Biobank. Source: &lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/37163049/&quot;&gt;Yun et al., 2023&lt;/a&gt;.
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;
&lt;p&gt;To compare disease-relevant polygenic risk scores (PRS) for different phenotypes, EDFs, and PPG or spirogram embeddings, a set of intermediate PRSes were first computed against each coordinate of the embeddings or pre-defined trait in EDFs. PRS for each coordinate can then be regressed against the target disease. Indeed, as shown in Figure 3, embedding-based PRSs can better stratify disease prevalence at different PRS percentiles. It seems a bit hard to interpret how much better the embedding-driven PRS is. Whether the enhanced stratification makes a meaningful difference in PRS depends on the heritability of each disease and whether we are sufficiently powered for the GWAS.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Optical Coherence Tomography autoencoder&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Concurrently, two other studies, &lt;a href=&quot;https://www.medrxiv.org/content/10.1101/2023.06.15.23291410v1&quot;&gt;Optical Coherence Tomogrpahy (OCT) autoencoder&lt;/a&gt; and &lt;a href=&quot;https://www.nature.com/articles/s42003-024-06096-7&quot;&gt;Unsupervised Deep Learning derived Imaging Phenotypes (UDIPs)&lt;/a&gt;, tried to use autoencoders instead of VAEs as the backbones to learn the embeddings. We’ll use OCT autoencoder to illustrate how autoencoder-based embedding might differ from VAE-based embedding.&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/rl_genetics/oct_overview.png&quot; /&gt;
    &lt;figcaption&gt;
        Figure 4: OCT image embedding learning. Source: &lt;a href=&quot;https://www.medrxiv.org/content/10.1101/2023.06.15.23291410v1&quot;&gt;Sergouniotis et al., 2023&lt;/a&gt;.
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;

&lt;p&gt;The OCT autoencoder paper also used data from the UK Biobank. OCT is a non-invasive imaging technique for the cross-sectional view of the human retina. Each OCT scan contains 128 cross-sectional images of the retina. A U-net was first trained on a set of 100 OCT scans that had manual segmentation maps to obtain an OCT thickness map generator. The retinal thickness maps of the left eye were used to obtain an embedding of 64 coordinates using the autoencoder (Figure 4). The OCT embedding was much larger than the PPG and spirogram embedding used in the REGLE paper (5-7 coordinates) to perhaps account for the increase in data volume.&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/rl_genetics/mtag.png&quot; /&gt;
    &lt;figcaption&gt;
        Figure 5: GWAS results for OCT autoencoder. Source: &lt;a href=&quot;https://www.medrxiv.org/content/10.1101/2023.06.15.23291410v1&quot;&gt;Sergouniotis et al., 2023&lt;/a&gt;.
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;

&lt;p&gt;Since the autoencoder-based embeddings were correlated, the authors performed GWAS on the embeddings, but also the first 25 principal components of the embeddings, and a multi-trait meta-analysis (&lt;a href=&quot;https://www.nature.com/articles/s41588-017-0009-4&quot;&gt;MTAG&lt;/a&gt;), a computationally efficient method to jointly analyze multiple related traits (Figure 5). In total, 239 lead loci were identified, 118 of which remained significant following Bonferroni correction. The authors reserved a subset of the UK Biobank participants for replication analysis. A total of 17 loci were replicated in the end, most of which were linked to the retinal layer thickness parameters. To further demonstrate the utility of the embeddings, the authors further used survival analysis to describe the predictive value of the embeddings and how the embeddings can be used for risk stratification of diseases.&lt;/p&gt;

&lt;h3 id=&quot;contrastive-learning&quot;&gt;Contrastive learning&lt;/h3&gt;
&lt;p&gt;Distinct from reconstruction, contrastive learning obtains the embeddings by learning representations that are invariant to simple transformations. The model aims to represent different views of an input in a similar way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Image-based genome-wide association study&lt;/strong&gt;&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/rl_genetics/igwas.png&quot; /&gt;
    &lt;figcaption&gt;
        Figure 6: iGWAS. Source: &lt;a href=&quot;https://www.medrxiv.org/content/10.1101/2022.05.26.22275626v3&quot;&gt;Ziqian et al., 2022&lt;/a&gt;.
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;

&lt;p&gt;The image-based genome-wide association study (&lt;a href=&quot;https://www.medrxiv.org/content/10.1101/2022.05.26.22275626v3&quot;&gt;iGWAS&lt;/a&gt;) obtained its embeddings for retinal fundus photos capturing structure information regarding the retina, optic disc and macula from the back of the eye.
The retinal fundus photos were first preprocessed into vessel segmentation masks. The embedding was trained on the segmentation masks using a modified version of &lt;a href=&quot;https://arxiv.org/abs/1801.07698&quot;&gt;ArcFace&lt;/a&gt;, an angular loss that maximises the distance between embeddings of different individuals while keeping the representations from the same individual close to each other (Figure 6.a). Specifically, the network was minimising the embedding distance between representations of the left and right retinas from the same individuals. Indeed, the cosine similarity is greater for matched retinas than random pairs in both the training dataset (EyePACS) and held-out test datasets (Messidor and UK Biobank) shown in Figure 6.b.&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/rl_genetics/corr.png&quot; /&gt;
    &lt;figcaption&gt;
        Figure 7: (a) upper right, embedding correlations; lower left, genetic correlations. (b) embedding correlations with retinal colors and 22 hand-crafted retina features. Source: &lt;a href=&quot;https://www.medrxiv.org/content/10.1101/2022.05.26.22275626v3&quot;&gt;Ziqian et al., 2022&lt;/a&gt;.
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;

&lt;p&gt;There are no benchmarks in comparing the embeddings produced by different machine learning methods. The embeddings obtained by contrastive learning have highly correlated clusters both on a phenotypical and genetic level (Figure 7.a). In particular, when assessing the embeddings with RGB values of the retina image, two clusters emerged. One cluster shows a high correlation with red and green values. And another cluster that has a greater correlation with the blue values. Given that the embeddings were designed to learn features related to vasculature, the structural arrangement of the retina, the embeddings should be invariant to the color of the retina. The authors admitted that better learning methods could alleviate the influence of retina color on the embedding space.&lt;/p&gt;

&lt;p&gt;Among all the novel genes being identified using the embedding GWAS, the authors performed a functional follow-up for a novel gene WNT7B. The WNT7B has only been known to be important in the blood-brain barrier development and its role in the retinal vessels has not been identified. To confirm the role of WNT7B gene, the authors compare the differences in mouse retinas in vivo by knocking off the Wnt7B gene using the short hairpin RNA technique which can silence target gene expression via RNA interference. It turns out that when WNT7B was knocked down, the total vessel area increased significantly in the intermediate vascular plexus but reduced in the deep vascular plexus. This functional follow-up provides the first experimental evidence to validate the biological effect of embedding-based genomic discovery.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cross-modal autoencoder&lt;/strong&gt;&lt;/p&gt;

&lt;div style=&quot;text-align: center;&quot;&gt;
    &lt;figure&gt;
    &lt;img src=&quot;/assets/images/rl_genetics/cvd_overview.jpg&quot; /&gt;
    &lt;figcaption&gt;
        Figure 8: Cross-modal cardiovascular state learning. Source: &lt;a href=&quot;https://www.nature.com/articles/s41467-023-38125-0&quot;&gt;Radhakrishnan et al., 2023&lt;/a&gt;.
    &lt;/figcaption&gt;
    &lt;/figure&gt;
&lt;/div&gt;
&lt;p&gt;The final method that I want to cover here is a cross-modal embedding for the cardiovascular state. The cardiovascular state embedding was trained using cardio MRI and electrocardiogram (ECG) data, two complementary data modalities about the human heart. cross-modal, I think the authors want to imply that the modalities are &lt;em&gt;paired&lt;/em&gt; and have &lt;em&gt;knowledge transfer&lt;/em&gt;. Cross-modal learning is multi-modal but multi-modal learning might not be cross-modal. Even though the embedding of this paper was not directly used in the GWAS input, I still want to talk about some of its method considerations:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;The embedding was evaluated on three downstream tasks including phenotype prediction, imputation and genomic discovery.&lt;/li&gt;
  &lt;li&gt;The embedding factors in information more than one modality.&lt;/li&gt;
  &lt;li&gt;In the embedding-based GWAS, the effect of confounders was removed using iterative nullspace projection, which reduces the dimensionality of the latent space that can be used to predict the confounders.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
  &lt;p&gt;Our results systematically integrate distinct diagnostic modalities into a common representation that better characterizes physiologic state&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When it comes to the design of embeddings, one could either develop a representation optimised for every downstream task. Or one could develop a universal representation to be used for different types of downstream tasks. Given that we do have a single physical state for our organs like how the heart is the physical manifestation of our cardiovascular state and more. Intuitively, we should aim to develop a common representation if we have sufficient measurements to scale up the impact of the embeddings.&lt;/p&gt;

&lt;p&gt;The objective function of the cross-modal embedding considers both reconstruction and contrastive loss as follows:&lt;/p&gt;

\[\begin{aligned}
&amp;amp;\mathcal{L}\left(\left\{X^{(j)} f_j, g_j\right\}\right)=L_{\text {Contrast }}\left(\left\{X^{(j)}, f_j\right\}\right)+\lambda L_{\text {Reconstruct }}\left(\left\{X^{(j)}, f_j, g_j\right\}\right),\\
&amp;amp;L_{\text {Reconstruct }}\left(\left\{X^{(j)} f_j, g_j\right\}\right)=\sum_{i=1}^n \sum_{j=1}^m\left\|x^{(i, j)}-g_j\left(f_j\left(x^{(i, j)}\right)\right)\right\|^2,\\
&amp;amp;\begin{aligned}
L_{\text {Contrast }}\left(\left\{X^{(j)} f_j\right\}\right)= &amp;amp; -\frac{1}{2} \sum_{I_k \in P_b} \sum_{j_1, j_2=1}^m \sum_{i=1}^{\left|j_k\right|} \log \left(\frac{\exp \left(e^{\text {temp }} f_{j_1}\left(x^{\left(i, j_1\right)}\right) \cdot f_{j_2}\left(x^{\left(i, j_2\right)}\right)\right)}{\sum_{i^{\prime}=1}^{\left|j_k\right|} \exp \left(e^{\text {temp }} f_{j_1}\left(x^{\left(i, j_1\right)}\right) \cdot f_{j_2}\left(x^{\left(i, j_2\right)}\right)\right)}\right) \\
&amp;amp; +\log \left(\frac{\exp \left(e^{\text {temp }} f_{j_1}\left(x^{\left(i, j_1\right)}\right) \cdot f_{j_2}\left(x^{\left(i, j_2\right)}\right)\right)}{\sum_{i^{\prime}=1}^{\left|j_k\right|} \exp \left(e^{\text {temp }} f_{j_1}\left(x^{\left(i, j_1\right)}\right) \cdot f_{j_2}\left(x^{\left(i, j_2\right)}\right)\right)}\right)
\end{aligned}
\end{aligned}\]

&lt;p&gt;Provided with input data with a subset of modalities \( X^{(i, j)}_{j\in\mathcal{I}} \), \( \mathcal{I} \in [m] \), where m is all the modalities available, we have an encoder \( f_j \) and a decoder \( g_j \).&lt;/p&gt;

&lt;p&gt;\( L_{\text {Contrast }} \) aims to reconstruct the samples. \(L_{\text {Contrast }}\) makes sure data points from the same modalities of the same participant are similar. \( \lambda \) is used to balance the importance between the reconstruction loss and the contrastive loss.&lt;/p&gt;

&lt;p&gt;Unlike previous studies for which the embeddings were directly used as the input for GWAS, the embeddings here were used to predict commonly used phenotypes such as the body-mass index and right ventricular ejection fraction to confirm it captures genotype-phenotype association for cardiovascular data. Not sure what the results might be if the GWAS was directly done on the embeddings.&lt;/p&gt;

&lt;h2 id=&quot;design-choices&quot;&gt;Design choices&lt;/h2&gt;

&lt;p&gt;It’s not easy to read through the related work in representation learning for genomic discovery because the existing works have used different modalities, evaluation metrics and machine learning methods (Table 2). What’s clear though is that by using embeddings, it will be possible to move away from labeled datasets, have less dependency on expert-curated features and identify novel loci that might not be discovered using conventional techniques.&lt;/p&gt;

&lt;p&gt;Also on a side note, I was amazed that the majority of the papers covered used data from the UK Biobank. Even though I work on the UK Biobank on a daily basis and so does everyone around me, it is incredible to see what people can do with rich resources like the UK Biobank on perhaps non-mainstream projects.&lt;/p&gt;

&lt;h4 id=&quot;table-2-characteristics-of-phenotype-representation-learning&quot;&gt;Table 2. Characteristics of phenotype representation learning&lt;/h4&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Method&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Learning objective&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Model Architecture&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Data source&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Embedding Size&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;GWAS hits&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;REGLE&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Reconstruction&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Variational autoencoder&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;170K PPG &lt;br /&gt; 351K spirograms&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;5&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;PPG: 40 known, 50 novel&lt;br /&gt;spirogram: 596 known, 63 novel&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;OCT autoencoder&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Reconstruction&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Autoencoder&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;31K  OCT images&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;64&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;118: 17 were replicated&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;UDIP&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Reconstruction&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Autoencoder&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;91K brain MRIs&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;256&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;199: 145 novel&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;iGWAS&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Contrastive&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;ConVNets&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;105K fundus images&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;128&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;34: 21 novel&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Cross-modal autoencoder&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Contrastive  &amp;amp; Reconstruction&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Autoencoder&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;45K cardic MRIs, 39K ECG&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;256&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;NA&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;I would want to highlight some of the key design choices that are important when developing the embedding for genomic discovery:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Evaluation metric&lt;/strong&gt;: Even though the end goal of the embedding, is to identify all the gene variants associated with a certain trait, Most of the embedding developed has not been evaluated for their usefulness for genomic discovery other than their training objectives. Metrics such as the heritability of an embedding, and relevance to diseases need to be assessed to develop the most relevant embedding (&lt;a href=&quot;https://www.google.com/search?client=safari&amp;amp;rls=en&amp;amp;q=EmbedGEM%3A+A+framework+to+evaluate+the+utility+of+embeddings+for+genetic+discovery&amp;amp;ie=UTF-8&amp;amp;oe=UTF-8&quot;&gt;EmbedGEM&lt;/a&gt;)&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Scaling laws&lt;/strong&gt;: By scaling laws, I am talking about the minimal amount of data and modal capacity needed to represent the state of our biology. For instance, in the cross-modal autoencoder paper, the network only had 10 million parameters to represent the cardiovascular state from cardiac MRIs and ECG data. If we are thinking about the complexity of the human heart, it seems a bit too small.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Embedding size&lt;/strong&gt;: Different modalities have different amounts of information. REGLE deliberately chose to smaller embedding space as the authors argued that it is better to have low-dimension uncorrelated embeddings and high-dimension correlated embeddings. Other approaches did not explicitly consider the influence of the embedding size w.r.t. the input information or the downstream GWAS.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Learning paradigm&lt;/strong&gt;: The machine learning techniques used thus far center around using autoencoder as the backbone with a reconstruction loss and contrastive loss. We don’t have any good data comparing the performance of different approaches.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Universal representation&lt;/strong&gt;: Ideally, we should be able to obtain a single encoder that generalises across populations. Having a universal representation will be computationally efficient as the users don’t need to obtain the embedding encoders on new datasets. Furthermore, the high-content clinical data measures some complex biology of the human body that has a universal representation in the real world.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;final-thoughts&quot;&gt;Final thoughts&lt;/h2&gt;
&lt;p&gt;We are still in the early days of understanding the representation learning of the high-content clinical data for genomic discovery. As our measurement techniques, we will inevitably acquire richer high-content data in large volumes. Leveraging data-driven approaches to understand complex data modalities might help us understand our biology in a way not possible before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Acknowledgment&lt;/strong&gt;: I would love to thank the following people I chatted with on the ideas related to this post: Karl Simth, Nina Cai, Chris Nellåker, Chris Yau, Alkes Price, Angus Burns, Xilin Jiang, Steven Lin and Laura Portas.&lt;/p&gt;
</description>
        <pubDate>Tue, 30 Apr 2024 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2024/04/30/RL_genetics.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2024/04/30/RL_genetics.html</guid>
        
        <category>representation learning</category>
        
        <category>Machine Learning</category>
        
        <category>genetics</category>
        
        
      </item>
    
      <item>
        <title>Why Human Activity Recognition Using Wearables Is Far From Being Solved</title>
        <description>&lt;p&gt;Authors: Hang Yuan &amp;amp; &lt;a href=&quot;https://www.bdi.ox.ac.uk/Team/rosemary-walmsley-1&quot;&gt;Rosemary Walmsley&lt;/a&gt; &amp;amp;  &lt;a href=&quot;https://scholar.google.co.uk/citations?user=-FqhzRcAAAAJ&amp;amp;hl=en&quot;&gt;Shing Chan&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Keywords: Human activity recognition, IMU, wearables, machine learning&lt;/p&gt;

&lt;h2 id=&quot;har-intro&quot;&gt;HAR intro&lt;/h2&gt;

&lt;p&gt;Human activity recognition (HAR) is a popular application for wearable devices. HAR describes the techniques that classify human activities from time series. In HAR applications, we often use data from several modalities, such as images and accelerometers/Inertial measurement units (IMUs). This post will focus on the issues related to the most common data modality, the IMUs.&lt;/p&gt;

&lt;p&gt;There are many examples of HAR applications in our daily lives: fitness and sleep quality tracking in smartwatches, human-computer interaction support in VR devices, and patient monitoring in clinical applications. On the surface, we might have the false perception that activity recognition from IMUs is already perfect and that nothing further needs to be done. Contrary to popular belief, among those who stopped using their wearable devices: 36% cited the perceived measurement in accuracy and 34% cited the incorrect activity tracking (&lt;a href=&quot;https://www.sciencedirect.com/science/article/pii/S0747563219303127&quot;&gt;Attig, C., &amp;amp; Franke, T. 2020&lt;/a&gt;). In fact, &lt;strong&gt;we’d argue that HAR is far from being solved&lt;/strong&gt; because of the following reasons:&lt;/p&gt;

&lt;p&gt;I. &lt;a href=&quot;#i-hard-to-define-what-is-an-activity&quot;&gt;Difficult to define what is an activity&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;II. &lt;a href=&quot;#ii-heterogeneous-benchmark-baselines&quot;&gt;Heterogeneous benchmark baselines&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;III. &lt;a href=&quot;#iii-getting-ground-truth-data-for-har-is-both-expensive-and-difficult&quot;&gt;Getting ground-truth data for HAR is both expensive and difficult&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;IV. &lt;a href=&quot;#iv-diverse-characteristics-lead-to-different-activity-profiles&quot;&gt;Diverse characteristics lead to different activity profiles&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;V. &lt;a href=&quot;#v-lacking-standarslization-in-data-storage-processing-and-analytics&quot;&gt;Lacking standardisation in data storage, processing and analytics&lt;/a&gt;&lt;/p&gt;

&lt;h2 id=&quot;i-hard-to-define-what-is-an-activity&quot;&gt;I. Hard to define what is an activity&lt;/h2&gt;
&lt;p&gt;For us humans, it is obvious when someone is running or doing dishes. However, it is much harder for machines to know what constitutes an activity. Take walking, an apparently simple behaviour, as an example. Despite its perceived simplicity, designing a step counter is non-trivial.&lt;/p&gt;

&lt;p&gt;Below you can find three different gait patterns:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Regular walk&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Irregular walk&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Edler Strolling&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;&lt;img src=&quot;/assets/gifs/walk1.gif&quot; width=&quot;300&quot; /&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;&lt;img src=&quot;/assets/gifs/walk2.gif&quot; width=&quot;300&quot; /&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;&lt;img src=&quot;/assets/gifs/walk3.gif&quot; width=&quot;300&quot; /&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Figure 1: Walking patterns. Source: GIPHY&lt;/p&gt;

&lt;p&gt;Depending on someone’s age and context, even a simple action like gait can come in many shapes.  A young adult might have a regular gait cycle. However, an elder might have a gait pattern that’s anything but regular. The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;duration&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;trajectory&lt;/code&gt;, and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;morphology&lt;/code&gt; of even the same activity type can be vastly different.&lt;/p&gt;

&lt;p&gt;Maybe it is hard to define exactly what is a gait. One might be tempted to define many gait subtypes to account for the differences in how people walk in different contexts. What a brilliant idea!  &lt;a href=&quot;https://sites.google.com/site/compendiumofphysicalactivities/home?authuser=0&quot;&gt;The Compendium of Physical Activities&lt;/a&gt; (Ainsworth, et al., 2000) is one of the major initiatives that aim to have a universal activity taxonomy. The compendium is widely used in epidemiological studies. More recently, &lt;a href=&quot;https://ego4d-data.org&quot;&gt;ego4d&lt;/a&gt; (Grauman, et al., 2022) also proposed something similar by having over 200+ activity labels to apply to its ego-centric video stream for VR. Depending on the application, we might choose a different activity dictionary.&lt;/p&gt;

&lt;p&gt;Nonetheless, it is important to note that none of the activity compendiums is perfect. In an ideal world, all we need is a single model that can classify every possible activity type. Unfortunately, we won’t be able to do that mainly because we will need a lot of data with annotated human activity. Right now, even the largest HAR dataset that we are aware of is too small to develop a model like that. More often, it suffices to develop a classifier using a much simpler definition. For example, if we only want a rough idea of how active someone is in general, we might be happy with a classifier that can separate sleep, sedentary behaviour, light physical activity, and moderate-to-vigorous physical activity like &lt;a href=&quot;https://bjsm.bmj.com/content/56/18/1008.abstract&quot;&gt;Walmsley, et al., 2022&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Another reason why HAR can be challenging is that we need an evolving activity definition to account for everything that we want to capture.  As new hardware surfaces, we need to have novel gesture recognition to improve the human-computer interaction (HCI) process. For example, Apple Watch’s battery doesn’t last very long, so the Watch tries to preserve the battery by keeping the screen dim until a specific uplift motion is detected. The lift detector is designed to capture when the user looks at the watch. This motion is unique and specific to the type of device. A general HAR model won’t be able to capture this type of motion, hence, adding more complexity to the HAR model development.&lt;/p&gt;

&lt;p&gt;Finally, something might be fundamentally wrong with how we treat an activity class. The current state-of-the-art HAR models often assign one activity class to a fixed-length window, e.g. 10 seconds. In a three-class classification where we try to discriminate between &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;walk&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;run&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sleep&lt;/code&gt;,  the model receives the same penalty in training regardless of whether the model misclassifies &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;walk&lt;/code&gt; as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sleep&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;run&lt;/code&gt;. The main reason for the undifferentiated penalty is that the model assumes that all the classes have an equal distance in the label space (Figure 2 left). The commonly used one-hot encoding label format exemplifies the equal-distance assumption. To mimic the true-distance label distribution (Figure 2 right), the model should receive a greater penalty if a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;walk&lt;/code&gt; sample is mistaken as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sleep&lt;/code&gt; than &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;run&lt;/code&gt;. Many studies have tried to convert the discrete input format into a continuous representation that reflects the data similarities as seen in language learning like word2vector. However, much less work has been done in the representation learning of the label space for HAR. To help resolve this issue, when annotating the data, we shouldn’t only just assign a window with one label. For ambiguous cases especially, we might benefit from noting down all the possible labels for a window. Then, we could employ techniques such as &lt;a href=&quot;https://ojs.aaai.org/index.php/HCOMP/article/view/21986&quot;&gt;soft label learning&lt;/a&gt; to enhance our model performance.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/act.png&quot; alt=&quot;Figure 2: Equal distance label space vs true distance labe space&quot; style=&quot;width: 100%; display:block; margin: 0 auto;&quot; /&gt;&lt;/p&gt;

&lt;h2 id=&quot;ii-heterogeneous-benchmark-baselines&quot;&gt;II. Heterogeneous benchmark baselines&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The benchmark datasets for HAR are so heterogeneous that as a field, we don’t know how to make an apple-to-apple comparison for different modelling techniques.&lt;/strong&gt; In popular machine learning conferences nowadays, there is a big emphasis on beating the &lt;em&gt;state-of-the-art&lt;/em&gt; performance on existing benchmarks. In the field of computer vision, for example, one can evaluate a new method on &lt;a href=&quot;https://www.image-net.org&quot;&gt;ImageNet&lt;/a&gt; or &lt;a href=&quot;https://cocodataset.org/#home&quot;&gt;COCO&lt;/a&gt;; if a paper proposes an algorithm that beats the current best model on these benchmark datasets, then that paper becomes the new &lt;em&gt;state-of-the-art&lt;/em&gt;. How much a contribution a paper largely depends on the performance difference between the proposed algorithm and the existing best method. When a well-recognized benchmark exists, it is easy to make an apple-to-apple comparison between different methods. However, in HAR, we don’t have a well-recognized benchmark yet, which makes it much harder to identify the method with the best performance.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/baseline_har.png&quot; alt=&quot;Source: Yuan et al. 2022 Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data &quot; style=&quot;width: 90%; display:block; margin: 0 auto;&quot; /&gt;&lt;/p&gt;

&lt;p&gt;There have been many open-source HAR benchmarks for researchers to use (Table 1). However, no existing dataset has become the gold standard for comparing different algorithms. The reasons are multi-faceted:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Most of the benchmarks are small&lt;/strong&gt;. For small datasets, cross-validation is better suited to provide a more robust estimation of the empirical risk as compared to the larger datasets. For evaluation on a larger dataset, usually, a subset is held out as the test set instead. Having a common test set makes it easy to compare different methods. In most HAR research, however, people rely on their own way of data partitioning, making it impossible to compare results from different papers directly.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;The limited sizes of the benchmarks also mean that the number of activity classes labelled is also limited.&lt;/strong&gt; Often, we will see almost perfect performance on some of the smaller datasets, but that doesn’t mean the method used is perfect for HAR. Especially, for small lab-based benchmarks, only a few activity classes are included, thus it is easier to obtain a great performance.&lt;/li&gt;
  &lt;li&gt;We define an activity label over a fixed window length. However, &lt;strong&gt;the current benchmarks have vastly different window length definitions in their evaluation, making it hard to even compare the model performance across different datasets.&lt;/strong&gt;&lt;/li&gt;
  &lt;li&gt;Lastly, &lt;strong&gt;data collected in a lab environment doesn’t truly reflect the model performance in the real world.&lt;/strong&gt; Admittedly, it is much easier to set up some mounted cameras in a lab so that we can label the data by looking at the video stream. However, people will likely behave differently in a lab and free-living environment. So to fully appreciate the performance of HAR, we need to test our model on more datasets collected under free-living conditions.&lt;/li&gt;
&lt;/ul&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Dataset&lt;/th&gt;
      &lt;th&gt;#Samples&lt;/th&gt;
      &lt;th&gt;Evaluation method&lt;/th&gt;
      &lt;th&gt;Window length&lt;/th&gt;
      &lt;th&gt;Evaluation metric&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Capture24&lt;/td&gt;
      &lt;td&gt;910K&lt;/td&gt;
      &lt;td&gt;Held-one-subject_out&lt;/td&gt;
      &lt;td&gt;30 sec&lt;/td&gt;
      &lt;td&gt;F-measure/Kappa&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Rowlands&lt;/td&gt;
      &lt;td&gt;36K&lt;/td&gt;
      &lt;td&gt;Tested proprietary algorithms with all subjects being in one test set.&lt;/td&gt;
      &lt;td&gt;1 min&lt;/td&gt;
      &lt;td&gt;ROC&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;WISDM&lt;/td&gt;
      &lt;td&gt;28K&lt;/td&gt;
      &lt;td&gt;10-fold CV&lt;/td&gt;
      &lt;td&gt;5 sec/10 sec&lt;/td&gt;
      &lt;td&gt;Accuracy&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;REALWORLD&lt;/td&gt;
      &lt;td&gt;12K&lt;/td&gt;
      &lt;td&gt;10-Fold CV&lt;/td&gt;
      &lt;td&gt;1 sec&lt;/td&gt;
      &lt;td&gt;F-measure&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Opportunity&lt;/td&gt;
      &lt;td&gt;3.9K&lt;/td&gt;
      &lt;td&gt;Fixed train/test split&lt;/td&gt;
      &lt;td&gt;500 ms&lt;/td&gt;
      &lt;td&gt;F-measure/AUC&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;PAMAP2&lt;/td&gt;
      &lt;td&gt;2.9K&lt;/td&gt;
      &lt;td&gt;9-fold CV&lt;/td&gt;
      &lt;td&gt;5.12 sec&lt;/td&gt;
      &lt;td&gt;F-score/Accuracy&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ADL&lt;/td&gt;
      &lt;td&gt;.6k&lt;/td&gt;
      &lt;td&gt;None specified&lt;/td&gt;
      &lt;td&gt;No fixed window length&lt;/td&gt;
      &lt;td&gt;Accuracy&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;As for the data-sampling variations shown above, from experience, it doesn’t make a big difference for analysis with deep learning models. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;30 Hz&lt;/code&gt; is usually a good threshold between battery consumption without the loss of performance for IMUs. Since most of the frequency content in human motion falls below &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;15 Hz&lt;/code&gt; (send us a reference if you have one!), going below &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;30 Hz&lt;/code&gt; will risk losing information useful for classification as suggested by &lt;a href=&quot;https://journals.humankinetics.com/view/journals/jmpb/4/4/article-p298.xml&quot;&gt;Small, et al., 2021&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;iii-getting-ground-truth-data-for-har-is-both-expensive-and-difficult&quot;&gt;III. Getting ground-truth data for HAR is both expensive and difficult&lt;/h2&gt;
&lt;p&gt;One of the key reasons why existing benchmark datasets is small is that it is challenging to annotate ground truth for IMUs. To annotate HAR datasets, we will require concurrent ACC and video data. The difficulties are:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;We sync the timestamps on both the wearable and video recording devices, for which the timestamps might not be in perfect synchrony.&lt;/li&gt;
  &lt;li&gt;It might be easy to obtain a video stream of human activity in a lab environment. However, data collected in a lab doesn’t reflect the data distribution in a free-living environment. However, getting the concurrent video stream in a free-living environment is much harder because we would require the participants to wear an ego-centric camera or install many cameras in the participants’ living environments. Neither of these is ideal, and both cause privacy concerns.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/cameras.png&quot; alt=&quot;Figure 3: Egocentric camera setup. Source: Ego4d and Capture24&quot; style=&quot;width: 60%; display:block; margin: 0 auto;&quot; /&gt;&lt;/p&gt;

&lt;h3 id=&quot;impossible-to-annotate-the-video-at-a-high-frame-rate&quot;&gt;Impossible to annotate the video at a high frame rate&lt;/h3&gt;
&lt;p&gt;Below is a list of images taken when I wore an ego-centric camera. An annotator would need to select an activity from 200+ activity classes for every picture taken. Depending on how many pictures are taken per second, the sheer volume of the task becomes extremely large. If one image is taken per second, then one will need to annotate 1 * 60 * 60 * 24 = 86400 images just for one day of data per person. We are no way near to having the capacity to have high-quality free-living data at the moment. The best dataset that we are aware of at the moment is the capture-24 which only takes one image every 30 seconds. On the other extreme, we have ego4d, which has a very high frame rate but a much shorter duration.&lt;/p&gt;

&lt;p&gt;Capture-24 and ego4d used different approaches to annotate human activity. The key difference between Capture-24 and ego4d lies in the sampling rate of the camera footage. Exactly how much signal is lost when doing the annotation with a lower sampling rate high depends on the sort of behaviour that we try to capture. We could also argue that activity variations over too short a time, for example, &amp;lt;1s or &amp;lt;5s, just shouldn’t represent a separate behaviour because humans don’t really shift behaviour that quickly. How quickly behaviour changes is also population specific. Activity transitions are likely to happen more quickly for kids than adults.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/sample_view.png&quot; alt=&quot;Figure 4: Egocentric view&quot; style=&quot;width: 80%; display:block; margin: 0 auto;&quot; /&gt;&lt;/p&gt;

&lt;h2 id=&quot;iv-diverse-characteristics-lead-to-different-activity-profiles&quot;&gt;IV. Diverse characteristics lead to different activity profiles&lt;/h2&gt;
&lt;p&gt;For the same type of activity, we shall expect to see large variations across populations. These activity variations are problematic because the model trained on a young healthy population is going to perform worse on an older population for example.&lt;/p&gt;

&lt;p&gt;Many aspects contribute to the variations we see in different population profiles:&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Age&lt;/li&gt;
  &lt;li&gt;Weight&lt;/li&gt;
  &lt;li&gt;Height&lt;/li&gt;
  &lt;li&gt;Arm length&lt;/li&gt;
  &lt;li&gt;Occupation&lt;/li&gt;
  &lt;li&gt;Education&lt;/li&gt;
  &lt;li&gt;Income level&lt;/li&gt;
  &lt;li&gt;Nationality&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The easiest solution is to collect more data, which we sadly don’t have most of the time, especially for labelled data. The alternative is to use methods that can make use of unlabelled data, which is much easier to get hold of. Relevant works include using transfer learning to personalise the prediction trained on a large pool of subjects to a specific subject for which we have very limited data.  Self-supervised-learning can also be used to learn useful embedding from unlabelled data such that the eventual model only needs to be fine-tuned on a much smaller labelled dataset. (&lt;a href=&quot;https://arxiv.org/abs/2102.06073&quot;&gt;Tang et al., 2021&lt;/a&gt;, &lt;a href=&quot;https://arxiv.org/abs/2202.12938&quot;&gt;Haresamudram et al., 2022&lt;/a&gt;, &lt;a href=&quot;https://arxiv.org/abs/2206.02909&quot;&gt;Yuan et al., 2022&lt;/a&gt;).&lt;/p&gt;

&lt;h2 id=&quot;v-lacking-standardisation-in-data-storage-processing-and-analytics&quot;&gt;V. Lacking standardisation in data storage, processing and analytics&lt;/h2&gt;
&lt;p&gt;Last but not least, &lt;strong&gt;HAR is hard not just because solving the problem itself is hard but also because wearable tech is still in an early stage so there is a lack of standardisation in data storage, processing and analytics.&lt;/strong&gt;  For the research-grade and consumer-grade smartwatches, because of the lack of standardisation, each vendor is using its own bespoke data storage format best suited for its own device. There are times when diversity is good, but different data formats bring even more differences in data processing and analytics pipelines. Even though they might try to capture the same data modality, we are still entirely sure whether the data collected and processed by different devices are comparable. Therefore,  many validation studies have been done just to know how to explain the device difference and how the device differences might contribute to contradicting conclusions in research studies (&lt;a href=&quot;https://bmcresnotes.biomedcentral.com/articles/10.1186/1756-0500-7-952&quot;&gt;Tully et al., 2014&lt;/a&gt;, &lt;a href=&quot;https://www.mdpi.com/1424-8220/22/16/6317&quot;&gt;Miller et al., 2022&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Understandably, it might not be easy to make a standardised format for everyone because each device manufacturer might have different data format requirements. However, we will save a lot by introducing these standards as a field. Having a universal standard would require buy-in from all state-holders, commercial companies, research labs and end-users to agree and design the best all-purpose future-proof solutions that everyone can benefit from. Although some effort has been made in the industry, such as the project &lt;a href=&quot;https://facebookresearch.github.io/Aria_data_tools/&quot;&gt;Aria&lt;/a&gt; developed by Meta, the field has not picked up momentum just yet.&lt;/p&gt;

&lt;h1 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h1&gt;
&lt;p&gt;In summary, the wearables field is rapidly growing, as exemplified by the recent launch of the Apple Watch ultra, which even incorporates a temperature sensor into its design. Maybe more data modality will enhance the applicability of wearables for HAR in the future. With the additional data modality, we can better capture the activities where movement is not the best indicator of activity intensity. For example, when using IMUs to classify weight lifting or construction work, the wrist movement will be slow, without concurrent physiological data, it will be hard to separate strenuous work from other light activities.&lt;/p&gt;

&lt;p&gt;There is still a lot of work to be done around the hardware before we can move towards a more unifying data storage and analytics standardisation. Equally importantly, we need to develop novel methods to understand the buried information and then potentially translate the information into more actionable insights.&lt;/p&gt;

&lt;h2 id=&quot;references&quot;&gt;References&lt;/h2&gt;
&lt;p&gt;[1] Ainsworth, B. E., Haskell, W. L., Whitt, M. C., Irwin, M. L., Swartz, A. M., Strath, S. J., … &amp;amp; Leon, A. S. (2000). &lt;a href=&quot;https://www.researchgate.net/profile/Ann-Swartz-2/publication/12330586_Compendium_of_Physical_Activities_an_Update_of_Activity_Codes_and_MET_Intensities/links/0912f51407bee1e3a6000000/Compendium-of-Physical-Activities-an-Update-of-Activity-Codes-and-MET-Intensities.pdf&quot;&gt;Compendium of physical activities: an update of activity codes and MET intensities. Medicine and science in sports and exercise&lt;/a&gt;, 32(9; SUPP/1), S498-S504.&lt;/p&gt;

&lt;p&gt;[2] Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., … &amp;amp; Malik, J. (2022). &lt;a href=&quot;https://openaccess.thecvf.com/content/CVPR2022/html/Grauman_Ego4D_Around_the_World_in_3000_Hours_of_Egocentric_Video_CVPR_2022_paper.html&quot;&gt;Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition&lt;/a&gt; (pp. 18995-19012).&lt;/p&gt;

&lt;p&gt;[3] Yuan, H., Chan, S., Creagh, A. P., Tong, C., Clifton, D. A., &amp;amp; Doherty, A. (2022). &lt;a href=&quot;https://arxiv.org/abs/2206.02909&quot;&gt;Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data&lt;/a&gt;. arXiv preprint arXiv:2206.02909.&lt;/p&gt;

&lt;p&gt;[4]Walmsley, R., Chan, S., Smith-Byrne, K., Ramakrishnan, R., Woodward, M., Rahimi, K., … &amp;amp; Doherty, A. (2022). &lt;a href=&quot;https://bjsm.bmj.com/content/56/18/1008.abstract&quot;&gt;Reallocation of time between device-measured movement behaviours and risk of incident cardiovascular disease&lt;/a&gt;. British journal of sports medicine, 56(18), 1008-1017.&lt;/p&gt;

&lt;p&gt;[5] Small, S., Khalid, S., Dhiman, P., Chan, S., Jackson, D., Doherty, A., &amp;amp; Price, A. (2021). &lt;a href=&quot;https://journals.humankinetics.com/view/journals/jmpb/4/4/article-p298.xml&quot;&gt;Impact of reduced sampling rate on accelerometer-based physical activity monitoring and machine learning activity classification&lt;/a&gt;. Journal for the Measurement of Physical Behaviour, 4(4), 298-310.
Chicago&lt;/p&gt;

&lt;p&gt;[6] Tang, C. I., Perez-Pozuelo, I., Spathis, D., Brage, S., Wareham, N., &amp;amp; Mascolo, C. (2021). &lt;a href=&quot;https://arxiv.org/abs/2102.06073&quot;&gt;Selfhar: Improving human activity recognition through self-training with unlabeled data&lt;/a&gt;. arXiv preprint arXiv:2102.06073.&lt;/p&gt;

&lt;p&gt;[7] Haresamudram, H., Essa, I., &amp;amp; Plötz, T. (2022). &lt;a href=&quot;https://arxiv.org/abs/2202.12938&quot;&gt;Assessing the State of Self-Supervised Human Activity Recognition using Wearables&lt;/a&gt;. arXiv preprint arXiv:2202.12938.&lt;/p&gt;

&lt;p&gt;[8] Miller, D. J., Sargent, C., &amp;amp; Roach, G. D. (2022). &lt;a href=&quot;https://www.mdpi.com/1424-8220/22/16/6317&quot;&gt;A Validation of Six Wearable Devices for Estimating Sleep, Heart Rate and Heart Rate Variability in Healthy Adults&lt;/a&gt; Sensors, 22(16), 6317.
Chicago&lt;/p&gt;

&lt;p&gt;[9] Tully, M. A., McBride, C., Heron, L., &amp;amp; Hunter, R. F. (2014). &lt;a href=&quot;https://bmcresnotes.biomedcentral.com/articles/10.1186/1756-0500-7-952&quot;&gt;The validation of Fitbit Zip™ physical activity monitor as a measure of free-living physical activity&lt;/a&gt; BMC research notes, 7(1), 1-5.&lt;/p&gt;

&lt;p&gt;[10] Attig, C., &amp;amp; Franke, T. (2020). &lt;a href=&quot;https://www.sciencedirect.com/science/article/pii/S0747563219303127&quot;&gt;Abandonment of personal quantification: A review and empirical study investigating reasons for wearable activity tracking attrition&lt;/a&gt; Computers in Human Behavior, 102, 223-237.
Chicago&lt;/p&gt;

&lt;p&gt;[11] Collins, K. M., Bhatt, U., &amp;amp; Weller, A. (2022, October). &lt;a href=&quot;https://ojs.aaai.org/index.php/HCOMP/article/view/21986&quot;&gt;Eliciting and learning with soft labels from every annotator&lt;/a&gt;. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing (Vol. 10, No. 1, pp. 40-52).&lt;/p&gt;
</description>
        <pubDate>Thu, 10 Nov 2022 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2022/11/10/why_human_activity_recognition_is_far_from_solved.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2022/11/10/why_human_activity_recognition_is_far_from_solved.html</guid>
        
        
      </item>
    
      <item>
        <title>Why I Took Off My Headphones</title>
        <description>&lt;h2 id=&quot;why-i-took-off-my-headphones&quot;&gt;Why I took off my headphones&lt;/h2&gt;
&lt;p&gt;The young generation likes to have their headphones on. Some do this simply for fun, while others do it for a better reason and more efficient use of their time. This can be achieved by listening to audiobooks or being on a conference call when we have extra bandwidth in addition to what we are already doing.&lt;/p&gt;

&lt;p&gt;I used to be one of those who have their headphones on all the time, in the gym or on the way to somewhere. I enjoyed listening to great minds sharing their life-long learning, and I thought I was getting ahead of others because I could make better use of my time. But I cannot be any more wrong.&lt;/p&gt;

&lt;p&gt;Indeed, whenever we waste time, i.e. not doing something productive, we feel guilty. Our human nature is telling us that we need to work hard to both give life meaning and to least survive physically, a.k.a having enough money to eat and drink. But sometimes, we just try to optimize our productivity too much that we lose sight of what is around us. Filling our fragmented time with audiobooks is one of the things that we could easily overdo that we forget to see what we could have seen.&lt;/p&gt;

&lt;p&gt;We can almost consider “wearing headphones” some sort of mechanism that we use to compensate for the less productive activities we do, commuting, doing dishes, waiting in line, etc. But is it really healthy? I argue that there is a limit to how much we should be wearing the headphones, as I was certainly doing it so much that I felt something was missing.&lt;/p&gt;

&lt;p&gt;There were probably three things that I really liked but went away when I had my “headphones” on: &lt;strong&gt;opportunity to observe&lt;/strong&gt;, &lt;strong&gt;short mental break&lt;/strong&gt;, and the &lt;strong&gt;appreciation of the journey&lt;/strong&gt;.&lt;/p&gt;

&lt;h3 id=&quot;opportunity-to-observe&quot;&gt;Opportunity to observe&lt;/h3&gt;
&lt;p&gt;Some people see things others don’t. We can also see things that our old selves couldn’t. To be better at seeing things. We need to practice. Practice makes perfect. The best opportunities for this are just lying around the corner. During any time of the day, if you really look closely, you will find things that you probably haven’t never seen before: the number of lights in your living room, the shape of a plant, or maybe a tree in your neighbor’s backyard. Observing is a very active process that allows you to see new things even at a place where you have been many times. It will also make it easier for you to pick up patterns that others don’t, not to mention that the whole process is super fun.&lt;/p&gt;

&lt;p&gt;When I had my headphones on. I simply couldn’t have enough bandwidth to both listen to whatever that is playing, practice observing and actually observe.&lt;/p&gt;

&lt;h3 id=&quot;short-mental-break&quot;&gt;Short mental break&lt;/h3&gt;
&lt;p&gt;We all know the importance of going on vacation regularly: it brings us internal peace and better prepares us for the challenges to come. Holidays usually happen after a long period of intense work. What I have failed to grasp was that it is equally important to take short mental breaks throughout the day.&lt;/p&gt;

&lt;p&gt;In fact, there are perhaps activities in our lives that could count towards short mental breaks. Some find cooking therapeutic, some might go for a run to zoom out, and some others might simply take a shower. These activities all share a common feature, and that’s letting the brain stop thinking. Nonetheless, I made the mistake of keeping my brain engaged if I found it in an idle mode.&lt;/p&gt;

&lt;p&gt;I thought I could have learnt more but in fact, without letting my mind go off track. My cognitive functions certainly didn’t perform as well as I’d like to. It was easier for me to feel like I couldn’t work anymore. I also had a sense of mental fullness as I could only act reactively to incoming information but was less capable of thinking proactively.&lt;/p&gt;

&lt;h3 id=&quot;appreciation-of-the-journey&quot;&gt;Appreciation of the journey&lt;/h3&gt;
&lt;p&gt;One final issue that I discovered when putting my headphones on is that I was losing the appreciation of the journey and became too goal-oriented. In a business environment, sometimes to deliver a product to the client in time, we have to prioritize getting the end-product ready. Something is better than nothing, right. People are much more motivated to get something out in the end than postpone the delivery time.&lt;/p&gt;

&lt;p&gt;But life is not just the finish line. It is also about how we get there, and there are sceneries that we might have missed in the journey. I certainly don’t want my life to be just about achieving one goal after another. I would prefer to enjoy the journey to see what’s there than just trying to reach what’s planned, i.e., getting a bigger paycheck or publishing more papers.&lt;/p&gt;

&lt;p&gt;The process of getting somewhere is equally if not more important than the end goal. This is particularly true in scientific discovery, where it is relatively easy to come up with a conclusion, but the conclusion is only believable if people also have trust in the way that it is derived follows certain rigour and logical reasoning.&lt;/p&gt;

&lt;p&gt;I am not sure if a similar analogy will also hold in a startup space. In a sense, many companies (Reddit, Sony, and Slack) eventually got big not only because they tried to reach one target after another. They all had to pivot at some point during their growth because something else they discovered was more promising.&lt;/p&gt;

&lt;p&gt;Don’t get me wrong. I still very much enjoy audiobooks and podcasts. I am just no longer the guy who is trying to become uber-productive. Sometimes, less is more. I definitely think that not wearing my headphones gives me more room to breathe and more time to think. I am curious if “having your headphones all the time” is a way to say that little time is left for exploration, but all one does is exploitation. Individuals can suffer from too much exploitation, and so do organizations.&lt;/p&gt;

</description>
        <pubDate>Wed, 02 Feb 2022 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2022/02/02/never_too_busy.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2022/02/02/never_too_busy.html</guid>
        
        
      </item>
    
      <item>
        <title>Why Diversity Matters?</title>
        <description>&lt;p&gt;Diversity oh well Diversity. Being born in a country that puts the education above all else, it is not surprising that the first time that I heard this word was at a college fair in my hometown, Guiyang, China. It was a fair about the Top 50 US Universities. I still remember vividly what the keynote speaker told the audience, “What makes America great is diversity.” At the time, while still being in high school, I couldn’t comprehend its importance. It felt like when my mother asked me to read Jane Eyre in primary school, I finished reading the book as it was written, but I lacked the life experience to truly understand Bronté’s love story.&lt;/p&gt;

&lt;p&gt;As I grew older and started to travel around different places, I began to realise diversity’s increasing importance not just in academics but also in other areas like business and governance. Everyone knows that diversity can bring something different to the table, might it be different genders, ethnicities, academics backgrounds or sexual orientations. People often assume that the whole point of having a diverse team is to be politically correct or to look great when you put your team photo on the website’s front-page. However, what people tend not to understand is the value of diversity itself.&lt;/p&gt;

&lt;p&gt;Like many expats, who have been constantly moving across borders and “forcefully” put to work with the international teams, we sometimes take the benefits from diversity for granted so much that I don’t know why it matters anymore. This is especially true for students from top schools. The students from these schools study in an international company and then go work for a multi-national firm. What we do need to understand is much of today’s world does not have the same kind of atmosphere. Maybe our precious ivory towers sometimes are sabotaging the very purpose for their existence.&lt;/p&gt;

&lt;h4 id=&quot;we-dont-know-what-we-dont-know&quot;&gt;We don’t know what we don’t know&lt;/h4&gt;
&lt;p&gt;The main reason why people often underestimate the value of diversity is that we don’t know what we don’t know. There are three kinds
of knowledge out there. The ones that we know, the ones that we know that we don’t know, and the ones that we don’t know that we don’t know. The very value of diversity is exactly why it is underestimated in the eyes of the be-holders. Recently, I was part of a STEM event that took place in Shanghai. Most of the participants are from middle to middle-upper households studying at top Uni outside China. Many of the guys loved to joke about other straight guys acting very gay like. I know that these people were all very nice and held no prejudice towards gays at all. But what was missing was that they didn’t understand LGBTQ+ communities well enough to know these kinds of jokes are very inappropriate and could potentially hurt LBGTQ+ individuals’ feelings in some very subtle way. I know this only because I am a member of the LGBTQ+ community and I would probably not have noticed about this behaviour’s potential negative implications if I were in their shoes. Sometimes, there are just things that you don’t know that you don’t know, and the only way to know is to let others tell you.&lt;/p&gt;

&lt;h4 id=&quot;the-need-to-increase-diversity&quot;&gt;The need to increase diversity&lt;/h4&gt;
&lt;p&gt;Recently, Goldman Saches enacted a policy saying, “No IPO if board lacks diversity.” I mean, if Goldman knows that it matters, it probably does matter right? But why do we need to increase minority presence all across the board?&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;It is hard enough to be the only one different in the room. If you are caucasian going to a meeting, imagine if all the other people in the meeting are Asian, African or something, whatever you do, people will notice. It is very intimidating to be in a situation like this.&lt;/li&gt;
  &lt;li&gt;It only gets harder when people don’t agree with you when you are a member of the minority group. So let’s say you are used to being the only person that has a different skin colour, but you still have to fight when other people don’t agree with your viewpoint. I had the experience one time at a Hackathon and there was this team doing a video game showcasing the life of children who suffer from Autism. The game aimed to simulate the experience of an Autism child protagonist. There was this one level in the game where you had to dodge all the balls others threw at you because they didn’t understand the fact that you have Autism. The game had many other levels that tried to simulate the adversarial encounters that the Autism children might face, but the issue was many of the levels were so descriptive and dark that they produce unintended side-effects. To put it a simple way, the game itself had many controversies. Hence, a few female judges raised the concerns that this sort of games needed to go through some form of ethics approval but most other male judges were adamant about the game arguing the game had a good intention and should not require any ethics/moral ground justification. Male judges far outnumbered female judges so by the time that everyone expressed his/her own opinion, all we could remember was that most people thought the game was OK. Whether or not the game itself had legitimate issues is another conversation but how the decision-making process is governed by the demographics makeup of the body did not reflect the views of the general population. That’s why we need to increase the proportion of minority groups to be able to see all different angels fairly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Don’t get me wrong. I also have friends who sometimes complain that nowadays it is more difficult for a white male because of the whole diversity talk. But is this true? We do have many program quota that is dedicated to minority groups, so this might come off to seem like many people get in not because of how good they are but what they are. This argument is precarious. First, most leaders in government or business are white male, and this doesn’t seem right when we consider the fact that half of the world’s population is female and a substantial portion of the population is not white. That’s why we need to have those specific programs to support minority groups. Furthermore, it is exactly because we don’t have a large enough women presence in some organizations/sectors that we fail to attract more of the similar candidates to the same areas.&lt;/p&gt;

&lt;p&gt;Ok, I have to admit that diversity is not a panacea to all the problems we have. Having diverse groups means that we are going to create more friction and sometimes this frication might slow down the process and destabilises in one’s organization. Perhaps sometimes it is OK to say men are better at something and women are better at other things. For example, most teachers in primary education are female, and the staff in the human resource department also tend to be female but people hardly complains about the lack of diversity there. I think the real deal-breaker is whether people can have a choice in the end. Being a teacher or a human resource professional is accessible to both men and women, so people choose to do them if they have a passion for either teaching or working with people. However, when we are talking about leadership, management, or other more critical positions people often don’t get to use what they want. This same principle also applies to other demographic minority groups. Only when everyone has a fair chance of doing the same thing, not because of what they are but because of what they can do, can we say we have done our job right.&lt;/p&gt;

</description>
        <pubDate>Sun, 09 Feb 2020 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2020/02/09/Why_diversity_matters.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2020/02/09/Why_diversity_matters.html</guid>
        
        
      </item>
    
      <item>
        <title>What Remains of Man After the Singularity?</title>
        <description>
</description>
        <pubDate>Sat, 16 Jun 2018 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2018/06/16/sigularity.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2018/06/16/sigularity.html</guid>
        
        
      </item>
    
      <item>
        <title>The Things that I Wish I Would Have Known in College -- Goodbye Jacobs and Hello EPFL</title>
        <description>&lt;p&gt;On June 11th of this year, one of my best friends Noma from Zimbabwe updated 
her Facebook profile into the picture below:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/grad.JPG&quot; alt=&quot;Jacobs Graduation&quot; style=&quot;width: 60%; display:block; margin: 0 auto;&quot; /&gt;
From left to right is me, Noma and Thato (another good friend of mine from Lesotho) wearing Jacobs Robe and scarf. The description that Noma put for the photo was “Thank you for being my two pillars of strength for the past three years,
As we say at Jacobs, this is not goodbye but it’s a see you later!!!” In case
you are wondering, yes we finally graduated.&lt;/p&gt;

&lt;p&gt;Till this day, I can still vividly remember the anxiety I experienced when I had 
only two weeks left to complete  my thesis, and still had to rerun
the model with a new set of parameters; the joy I felt, when my family and
my friends’ families all came to our Jacobs from different distanced worlds for the first time and maybe for
the only time; the sentiment I went through, when my family and my friends said Goodbye to me on the same day, and I had to move to 
Tuebingen to start my summer internship at Max Planck Institute. Without a doubt, the ups and downs 
at the graduation season will stay with me for long. Bearing all these emotions, I intent to write down a few most important lessons that I learned from college to let the upcoming college students 
know what to look out for, and to remind myself of these valuable things which shall not be 
forgotten in the coming days.&lt;/p&gt;

&lt;h3 id=&quot;hard-work-will-be-paid-off-eventually&quot;&gt;Hard work will be paid off eventually&lt;/h3&gt;
&lt;p&gt;Going to an American Christian high school in Los Angeles is more likely to make me a better person than
a better mathematician or scientist, as the school curriculum will be a lot more focused on one’s spiritual 
and mental growth than one’s academic Al pursuits, and thus my first year in college was rather challenging for me. When I finally got to compete with people like those who are from Eastern European countries and who have already specialized their areas of studies during high school, I was intimidated. Remember the guy who can almost answer the questions that the teacher raises 
without much thinking in class? In high school, I was one of those annoying kids but when I started out college, I could no longer be the one of those not because I didn’t want to but because I couldn’t. In a sudden, I felt depressed, pessimistic and less confident.&lt;/p&gt;

&lt;p&gt;I had my first mental breakdown already during the first week of college, when I spent the whole weekend 
trying to finish a recursive proof from the general computer science lecture to no avail. Back in my mind,
I already thought about all the alternatives if I couldn’t finish my study program. Maybe I should simply drop out of school or maybe I should switch my major from computer science to something else, however, non of those alternatives made sense to me, so the only way out was to graduate and to keep on doing what I was doing. Fortunately, on the third day, after going through many sample proofs, I was finally able to write down a proof sketch for that homework and it was the time that made me realize that college would be significantly harder than high school if I wanted to do well.&lt;/p&gt;

&lt;p&gt;Thankfully when I was struggling and was suffocating myself, I met Noma who was having many similar classes as I did (she studied computer engineering and 
I studied computer science), and she constantly reminded me of the fact that, “if other people can do it, then we can do it,” and, “since those people
are already ahead, and if we want to win the competition in the future, we will have to work harder to make up the gap.” 
At that time, although we needed to spend a lot more time digesting the course materials, but we told ourselves that we should
feel lucky because we were learning more than the other people who had already known the course content. Gradually, we got used
to the heavy workload in college and kept on working harder. It might be due that the people who thought they knew everything were slack off in the beginning, and when the professors were teaching the new concepts that they didn’t know, they still tried to persuade themselves that they probably have seen the materials somewhere instead of starting anew like what we did. It was always a beginner’s mindset for us no matter whether or not we thought we knew about something. By the end of college, I could already deal with course load quite easily and achieved excellent marks. Both Noma and I did an exchange semester to enrich our learning experiences. She did hers at University of Pennsylvania and I did mine at Carnegie Mellon University. All of these results didn’t come from how smart we were but because the hard work that we put in. Being smart or having learned something doesn’t imply that one is going to learn more or do better in the future. The future rewards for what you have done and you are going to do but not how SMART you are. There are many studies done that support this claim from cognitive psychology or neuroscience but I will stop here. If you are interested in knowling more, you can read &lt;em&gt;The Talent Code&lt;/em&gt; by Daniel Coyle and &lt;em&gt;The Outliers&lt;/em&gt; by Malcolm Gladwell who popularized the 1000-hour rule.&lt;/p&gt;

&lt;h3 id=&quot;build-up-your-own-support-network&quot;&gt;Build up your own support network&lt;/h3&gt;
&lt;p&gt;Many overachievers in high school/college think they don’t need other people’s help in order to get what they want but this perspective probably should change. Confucius once said, “To be established is to help others established, to achieve is to help others achieve.” Having your own network for either social, mental, or academic purposes is a great way for people to help each other out. I have Noma and Thato as my two great companions along the way. They tell me how to not make stupid decisions when I get excited, they 
comfort my feelings when misfortunates happen, they cheer me up when I encounter hurdles and they support me to the very end when I am determined to do something with good reasons. I don’t know how would I have survived without these two kindest human beings that I know. I am sure that you know many other amazing people around you. Go to talk to them and include them as part of your tribe. Go to parties with them and give them the support that they need and you will get yours when the time comes.&lt;/p&gt;

&lt;p&gt;One might say, for feelings, and emotions having a support network is indeed great but a support group might not work for learning purposes. I would argue for the contrary. Learning in a small groups is extremely useful if you know how to cooperate well. I think a study group has become even more important in higher level courses, where only a few people know about the materials and your classmates will be the only ones that you can talk to.&lt;/p&gt;

&lt;p&gt;Often when we have some homework sets, we will always do our own problem sets first and than come for a discussions. In college, the grading might not come immediately after you submit the homework so you don’t know if you have misunderstood something, and then you carry the misunderstanding for a few weeks and even when the homework reviews are posted, you overlook your own misunderstanding and therefore you never get corrected. But by learning in a group after each individual has put enough effort of self-learning on his own, the discussion sessions can help to clarify the things that are confusing and are good to think about those abstract concepts from different perspectives. Sometimes when you have a disagreement with someone, it means that either someone is wrong or both of you are wrong and that’s a great learning opportunity that you won’t have by working alone. Lastly, even if you are really good student, explaining things to the others can only strengthen your own understanding because you will need to break down the difficult concepts into coherent and simple languages that anyone can understand. By going through those thought processes, you will gain a lot more about the materials for sure. Even the Nobel laureate Richard Feynman uses this so called Feynman technique to learn almost everything by explaining things to the others.&lt;/p&gt;

&lt;h3 id=&quot;make-good-use-of-your-summer-time&quot;&gt;Make good use of your summer time&lt;/h3&gt;
&lt;p&gt;Summer breaks give college students another great opportunity to advance our career and explore more places. Many people don’t realize how important their summers are until they start to look for a job and then they realize all positions require prior experiences in project management, software engineering or business development. They then get anxious and confused about how they can get experiences even when the entry level positions require prior experiences. The secret recipe is to do internships over the summer. It is true that many won’t land in their dream work place during their first internship, but that’s completely OK. One can start by working for an OK company during the first summer, then during the second year, because of your experience in the first company, you will be able to land in better places and hopefully you will get into your dream workplace by the time you graduate.&lt;/p&gt;

&lt;p&gt;Doing internships is great for a couple of reasons:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;It typically lasts 8-12 weeks, short enough to let you don’t suffer too much if you don’t like the job and long enough 
to let you get a taste of how it feels like working in certain company.&lt;/li&gt;
  &lt;li&gt;You can travel quite a bit at a low cost. You don’t always have to apply to the company of interest in local area. For instance, if want to work for Microsoft, Microsoft has many offices around the world. If you feel adventurous enough, just apply for the positions in Asia or in Europe, and you will be able to travel for free. Sometimes some countries have strict work regulations for foreigners but for an internship sorting out the paperwork shouldn’t be a big issue for major international cooperates.&lt;/li&gt;
  &lt;li&gt;For those who want to pursuit a research career, having internships at the labs you want to join is first going to help you find out if you and the lab are a good match and second increase your admission chance. For myself, I spent every single summer doing research at different places in my undergraduate career at my own university, in London and at Max Planck. Although they can seemly be similar on paper, the work environment, the supervision style and the city can change your happiness a lot. Doing several summer research jobs is even more important , if a phd is on your agenda. I am sure that you don’t want to commit 4-6 years of your life living in a place that you won’t be happy and the best way to make sure that doesn’t happen is to do an internship in advance.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3 id=&quot;learn-to-live-like-a-third-culture-kid&quot;&gt;Learn to live like a third-culture kid&lt;/h3&gt;
&lt;p&gt;I’ve come to probably the most important lesson that I learned in my whole college career and that is how to live like a third-culture-kid(TCK). TCKs refer to the group of individuals who have lived in other cultures other than the ones where they were born. I consider myself as a TCK who was fortunate enough to have traveled to so many countries and have studied in China, Germany, the US, the UK and Switzerland thanks for the generous support from my family. The number of people like me is growing as the world is becoming increasingly more globalized. This trend is very easy to see merely by looking the number of international and exchange students out there in different countries (Sorry Britain this doesn’t apply to you until you sort out your own politics). There are lots of benefits of being a TCK such as being able to speak multiple languages, having abundant travel experiences and being able to quickly adapt to a new environment. Indeed those experiences make us more mature and enrich our lives a lot but on the other hand, being a TCK brings many issues that many don’t know how to deal with.&lt;/p&gt;

&lt;p&gt;I think the most critical question that TCKs have is where do we actually belong. It is often the case when we move to another country from our homeland, we become nostalgic about the food, the culture and the people back home, nevertheless when we actually return, then we want to leave again because our brain just beautifies everything in our memory and the reality actually has many pitfalls that we seem to have forgotten. What’s more is even after moving out of the home country again, the foreign place doesn’t feel like home either, cultural and/or languages differences. So many people are in this loop of constant moving but don’t feel belong to anywhere at all. To me, I look at this issue by thinking slightly differently, that I do recognize that there is this physical home, Guiyang, China where I grew up and call home, even though I don’t really think I belong there and there is this home that’s always on the way, this home where other TCKs gather and can understand each other. It doesn’t have to be something concrete that one can touch but more of a concept that wherever we are, wherever is home.&lt;/p&gt;

&lt;p&gt;I can keep talking about TCK for pages, but really I just want to raise people’s awareness of the existence of this group of people who are becoming more prevalent nowadays. Note that although many people might have not lived in different countries for an extended amount of time, but also feel the same like there is a huge difference between the culture in a small village in Texas than the culture in New York, we call this group of people trans-culture-kids (TRCKs) which more or less share similar feelings and issues as the TCKs. To find out more about TCKs, David Pollock’s &lt;em&gt;Third Culture Kids: Growing Up Among Worlds&lt;/em&gt; can be a fascinating read. He talks about the benefits and challenges of being a TCK and some solution to counteract the major hurdles.&lt;/p&gt;

&lt;p&gt;There is one more thing that I’d like to address before moving to the next section is that TCKs are just like any other kid at the same time. It is common for TCKs to have a sense of superiority because of their colorful life experiences and they would tend to think other people as boring as they probably have just been living in the same city for their entire life. The truth is we can learn something from everyone. An ancient Greek wise man once said that, there is no one so that stupid that he can learn nothing from, and Confucius also has said, “When I walk along with two others, they may serve as my teachers.” So we shouldn’t be too proud of our privilege, should we?&lt;/p&gt;

&lt;h3 id=&quot;remember-to-have-fun&quot;&gt;Remember to have fun&lt;/h3&gt;
&lt;p&gt;Lastly, you should learn to enjoy the process as much as you can. We aren’t robots, and we have feelings. Indeed hard work pays well but we can’t be constantly living under stress. Our will power is limited, if you have to endure the boredom you experience doing things you don’t like, how do you have the will power to do the things that are more important? Eventually you are going to burn out. So really, remember to do the things you really enjoy doing and go for the opportunies you really like. Many people are concerned about losing faces or don’t dare. To be honest in 15 years, probably no one is going to remember that stupid question you asked in class or that awkward date you had with the person you had a crush on. But we, ourselves are going to regret more for the things that we didn’t do than the ones that we actually did.&lt;/p&gt;

&lt;p&gt;Again let me summarize what I found important in my college time:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Hard work will be paid off eventually.&lt;/li&gt;
  &lt;li&gt;Build up your own support network&lt;/li&gt;
  &lt;li&gt;Make good use of your summer time&lt;/li&gt;
  &lt;li&gt;Learn to live like a third-culture kid&lt;/li&gt;
  &lt;li&gt;Remember to have fun&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/lake.JPG&quot; alt=&quot;The Geneva Lake&quot; style=&quot;width: 80%; display:block; margin: 0 auto;&quot; /&gt;
The lake view in front of EPFL campus&lt;/p&gt;

&lt;h3 id=&quot;whats-next-at-epfl&quot;&gt;What’s next at EPFL&lt;/h3&gt;
&lt;p&gt;I really hope that you enjoyed reading this post and if you have experienced anything common in your college life. Please feel free
to comment below, I’d love to know the life lessons you learned in college and if other people experienced similar issues as I did. So that was the end of my undergrad career and I am already at EPFL doing the intensive French course before all the courses start. Till this day, I have been enjoying the wonderful weather and the beautiful lakes. EPFL even has a really good Chinese restaurant on Campus. What can I complain!?s I feel really excited to embark my next journey at EPFL and also my lab assistantship at professor Courtine’s lab on spinal cord. Keep in touch and more is to come :D&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/chinese.JPG&quot; alt=&quot;Chinese food at EPFL&quot; style=&quot;width: 60%; display:block; margin: 0 auto;&quot; /&gt;
The Chinese restaurant in Rolex Learning Center is AWESOME :D - said by a Chinese&lt;/p&gt;
</description>
        <pubDate>Mon, 28 Aug 2017 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2017/08/28/jacobs.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2017/08/28/jacobs.html</guid>
        
        
      </item>
    
      <item>
        <title>My First Date with Interdisciplinary College (IK)</title>
        <description>&lt;p&gt;Last week, I got the chance to attend &lt;a href=&quot;https://interdisciplinary-college.de&quot;&gt;IK&lt;/a&gt; at Guenne in Northern Germany. The whole week was packed with inspirational lecture series and thought-provoking conversations. I got to meet many scientists and students coming from different backgrounds that were interested in the intersection of brain sciences, life sciences, and artificial intelligence. I would totally love to go back to IK again next year if I get the chance.&lt;/p&gt;

&lt;p&gt;So you might be wondering what IK is. IK (Interdisciplinary College) as its name implies is a spring school that started 30 years ago which aims to promote the education and awareness of artificial intelligence from the perspectives of different disciplines. Participants around the world gather at the beautiful Heinrich-Luebke-Haus for a week, learn about the cutting-edge research, and discuss their understandings on the raised issues. This year IK keeps up its multidisciplinary nature by having talks from deep reinforcement learning to biological spiking neural networks, and from music theory at the cognitive level to the connections between Artificial Intelligence and sci-fi movies. What I appreciate more were the intensive discussions that I had with the scientists and my peers. I’ve never encountered a group of people that are so open about ideas and willingly share their thoughts with others. I felt overwhelmingly welcomed already even on the train to Soest where I met another IK participant. The discussions just kept coming. Although the participants were mostly German, I almost always felt not being left out in any talks I had as people would either only speak in English or switch to English when they noticed my poor German skills.&lt;/p&gt;

&lt;p&gt;I first knew about IK a year ago in the machine learning reading group with professor Herbert Jaeger, but it didn’t catch my eyes that much because IK was held during school time. If I wanted to attend, I would have to miss a whole week of school. It wasn’t until HPI Berlin Hackathon which occured in the summer of 2016, that I decided I would attend IK 2017. At the Hackathon I met Marvin Fey who I teamed up with for a Wiki-Data data analytics project. He told me his thoughts about intelligence which mainly came from Jeff Hawkins’ book On Intelligence and was telling me his fascinations of wanting to understand more about intelligence in general. I was simply stunned by his hours-long talk and was curious about how he found out his interests. Then he told me IK was the evolutional event for him. Immediately I figured: IK was the black swarm, the rare but impactful event that often changes people’s life trajectory. Because of this encounter, I made up my mind to become part of IK 2017 and while I am writing this blog post, I still think coming to IK was one of best decisions that I’ve made.&lt;/p&gt;

&lt;p&gt;Throughout the whole week, the days were packed with different lecture series and the nights were packed with some evening lectures and fun activities for social networking. I would say I had pretty good life-work balance, having roughly eight hours of lectures, eight hours of sleep and the rest for networking and discussions. There was this Epistemology (theory of knowledge and understanding) lecture given by Gerhard Roth from Uni Bremen. I would never have thought about the world from his viewpoint if I didn’t attend IK. In essence, he argues that because:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;the brain has no access to the outer world&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;the sensory organs transduce stimuli from the external environment into neural signals that affect brain states, however, those brain stats are meaningless themselves&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;furthermore the mind is mere a construct of the world&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;and therefore the brain cannot prove the “truth” of its own construct. This theory might sound insane from the first glance, but if one takes a closer look, the virtue will disclose itself. As humans, we are limited by our sensory inputs, our eyes can only observe lights within certain frequency band and so do our ears to sounds. We never see the world in its complete form and how do we really know the world we perceive is the real one? The answer is we don’t know, period. I am not sure if I am taking the view too far, but if we interpolate epistemology more, we can even say the sciences that are considered true, are only coherent with the world that we perceive. If coherence is reached among many people, then we will conclude that this scientific fact is true but in fact no one knows the real truth. I am not to judge the validity of epistemology but to give you a taste of what IK lets me ponder greatly. There are many other fascinating lectures such as understanding the evolution of society via network theory and robots like me by Ipke Wachsmuth who presented himself as a robot and opened up intense discussions on the ethics of robots in this fast-growing society. This list can go on and on, but I guess you already have an idea of what the lectures are like.&lt;/p&gt;

&lt;p&gt;As inspiring as the lectures might be, I still think listening itself is only a passive process. It’s hard to convert everything you hear into the language of your own. That’s why I like to facilitate my learning process using more active means i.e. talking to others and writing. During last week, I felt like I was always caught in some kind of discussions. The fun part was random people would simply join if they overheard the conversation and thought it was interesting. I want to share two messages that I learned over those discussions, and I hope you will like them too:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;People don’t realize they actually refer to different things when talking to each other, even when they are using the most common words like love, feelings and emotions. It’s just those words appear so often so that we have this misconception that we know their meanings when we really don’t. This can lead to many implicit misunderstandings as we tend to think other people are thinking like us, and thus assign the same kind of attributes when encountering those words. Some may say it is hard to define those abstract concepts and is it really necessary to reach an agreement? I think it is important because too many people have avoided answering those questions and just overlook the definitions of those common terms and words in general, but that’s not an excuse to not pursue this line of inquiries.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;Try not be judgmental in things you do, no matter if they are good or bad. The core message stemmed from Timothy Gallwey’s The Inner Game of Tennis. When we are playing tennis, we tend to be over-analytical. Although you might know you tend to raise your arms a little too high for backstrokes, you will find it hard to connect that concept with your body. That’s why you often hear people saying, “my body is not listening to me.” Words by themselves are meaningless. Think of this as if when someone is teaching you how to play tennis, he will first translate the way he plays into words, and you will need to digest those words into actions. The words are just an interface between the way that you play and the way that he plays after all. If you can just observe what he does, and try to imitate that, wouldn’t the training become more effective without the potential loss of information in between (words)? What one should do is to focus on the present and to feel the flow of his own body. The reason why one shall not judge is because when you judge, you start to lose focus, you start of think about, “Am I going too fast?”, “Are my arms too high up again?”, “I am not having a good day” and “I am just a clumsy player.” Even positive judgments can be harmful, because once you have been complimented by someone, you will want to live up to that expectation and when you do not, you start panicking. To be fully immersed in the process, one should not judge and one should go zen. Only when your attention is on the process, you start to notice the things that you rarely do and the things that words cannot describe, and that’s when one improves internally. But also remember that not judging doesn’t mean ignoring the mistakes but to take an impartial view on things that you do. This thought process can be applied in tennis, other sports, and life as well.&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;IK seemed to have empowered me with the curiosity about every aspect of sciences and life. I know many participants who I’ve talked to have the same feeling. Life simply rejuvenates. It is intellectually stimulating to be surrounded by computer scientists, cognitive scientists, neuroscientists and other scholars all at once. I am glad that I came to IK to get to know those amazing people at a personal level. I do hope I get to meet them again in the future to talk about their progress on research. Moreover, my IK experience was especially enriched by Marvin Fey. He seems to be a mentor and a good friend at the same time, guiding me through all the traditions the IK has and inspires me with his constant thought flow. We had lots of fun together. He also writes about life and brain sciences on his &lt;a href=&quot;http://www.velarys.com&quot;&gt;blog&lt;/a&gt;. He likes to explain abstract concepts in simple terms with precision and conciseness. If you happen to like my writing, you will enjoy his too : )&lt;/p&gt;

&lt;h2&gt;&lt;img src=&quot;/assets/images/ik2017.png&quot; alt=&quot;&quot; /&gt;&lt;/h2&gt;
&lt;p&gt;IK 2017 poster: designed by &lt;a href=&quot;http://mgerstenkorn.jimdo.com/&quot;&gt;Maike Gerstenkorn&lt;/a&gt;&lt;/p&gt;
</description>
        <pubDate>Mon, 20 Mar 2017 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2017/03/20/ik.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2017/03/20/ik.html</guid>
        
        
      </item>
    
      <item>
        <title>Undergraduate Research 101</title>
        <description>&lt;p&gt;Inspired by a friend of mine who just started his computer science career 
at CMU out of high school, I have finally found some time to write this long waited Undergraduate Research
101 to help more young and bright students to get to the right track to becoming future researchers. Normally
when you see a blog post like this, it is written from the perspective of the faculty members but this 
time I would like to give away  my five cents to my peers who are eager to get involved with cutting-edge 
technology from a student’s perspective.&lt;/p&gt;

&lt;p&gt;Similar to my other blog posts, I would like to structure the article into the answers to the following questions,
the reason I do this is that I believe having a question to think about while reading something engages oneself more
as compared to purely passive knowledge intake by glimpsing through lines of text:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;What is research?&lt;/li&gt;
  &lt;li&gt;Is research for me?&lt;/li&gt;
  &lt;li&gt;When should I start?&lt;/li&gt;
  &lt;li&gt;How should I start?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3 id=&quot;what-is-research&quot;&gt;What is research?&lt;/h3&gt;
&lt;p&gt;For this, I would highly encourage you to read about &lt;a href=&quot;http://matt.might.net/articles/phd-school-in-pictures/&quot;&gt;The illustrated guide to a Ph.D.&lt;/a&gt; written by &lt;a href=&quot;http://matt.might.net/&quot;&gt;Matt Might&lt;/a&gt; who is a professor of computer science at University of Utah. In figure I, Might invites us to think of the knowledge that we obtain after the completion of different school levels as circles. As we climb up the academia ladder, the circle that we leave behind will grow bigger and bigger and there exists this outer loop called human knowledge boundary as shown in figure II. The purpose of research is to try to make a bump out of that boundary at a narrow scope.
&lt;img src=&quot;/assets/images/research1.png&quot; alt=&quot;&quot; /&gt;
Figure I&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/assets/images/research2.png&quot; alt=&quot;&quot; /&gt;
Figure II&lt;/p&gt;

&lt;p&gt;From having a well-rounded education background to be able to do something that no one else has done before takes years of commitment, and yet starting from you undergraduate years will help you reach the outer layer faster. However as interesting as research might be, you need to know your priorities well when things get tough. Without a doubt, doing well in coursework deserves to be on the top of the list if not the most important one. The rationale is simple. Even if you are excited about how to build driveless cars or develop the next-generation super computers, you need to have solid fundaments in math, computer science and even writing so that you will be able to understand other people’s work and implement a new algorithm that you just came up with to see if it really works or not. Of course paying attention to courses in other disciplines might help you develop soft skills for future collaboration.&lt;/p&gt;

&lt;p&gt;Meanwhile, you need to understand doing research is different from taking classes. Doing well in class is important but it is by no means the ticket to become the next Albert Einstein for they require different kinds of skill sets. Conducting research often needs you to be creative in ways that no one else has ever been and work independently when your research question leaves out lots of uncertainty for you to figure out whereas taking classes require you to comprehend the concepts that have been fairly well studied and working on the problem sets under clear requirements.&lt;/p&gt;

&lt;p&gt;My sincere suggestion on course work is trying to do reasonably well in them and spend the rest of your time on other things that are also important such as research, sports, friends and family.&lt;/p&gt;

&lt;h3 id=&quot;is-research-for-me&quot;&gt;Is research for me?&lt;/h3&gt;
&lt;p&gt;Really the best way to find out is to try to find a professor in your university and work with him/her for a semester or two. If you are not sure if you even want to get started, you ask yourself a few questions to determine whether you will enjoy research or not.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Are you OK with working independently?&lt;/li&gt;
  &lt;li&gt;Are you OK with working on something of which the requirements are largely unspecified and you will have to find them out yourself?&lt;/li&gt;
  &lt;li&gt;Are you OK with getting stuck on something and you may not be able to find the solution until two weeks/months later?&lt;/li&gt;
  &lt;li&gt;Are you OK with not being able to find the answers on Google and still not freak out : )&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If all of your answers are yes or maybe yes, then I think you can give it a try. Otherwise, you would probably want to take more classes and explore different things before you dive into research.&lt;/p&gt;

&lt;h3 id=&quot;when-should-i-start&quot;&gt;When should I start?&lt;/h3&gt;
&lt;p&gt;I would say it is never too early to start. I started working with professor Kohlhase at Jacobs University towards the end of the first semester in college on &lt;a href=&quot;https://github.com/KWARC/sTeX&quot;&gt;Semantic Tex(sTex)&lt;/a&gt; and at that time I didn’t even know how to do git properly. Even if you don’t that is OK. You can still learn those things on the fly as long as you are motivated.&lt;/p&gt;

&lt;p&gt;For students who are from large research universities like CMU, it is easier to find world-class researchers who need people in their labs, but for those who aren’t in similar situations, you are not completely hopeless, there might still be opportunities in the neighboring institutes and maybe spending the summer at some good labs will also be a good choice.&lt;/p&gt;

&lt;p&gt;Regardless how much you have learned when you first enter college, by working in a research group even on some of the very trivial problems will help you gain a better perspective about the field that you are interested in and you will be able to appreciate more about what you study in lectures. For an example, I had formal language and logic (FLL) in my second year. FLL was a theory course which first introduced formal languages and many people were not motivated and lost interests when going through the tedious proofs. On the contrary, I had more incentives to do well because for the research that I was involved in used the knowledge of formal languages. It was the OMdoc and &lt;a href=&quot;http://www.openmath.org/standard/om20-2004-06-30/omstd20html-3.xml&quot;&gt;OpenMath schemata&lt;/a&gt; that I needed to look into which properly specified the the markups for mathematics. I knew the reasons behind studying the different kinds of languages and grammars and I was applying those knowledge into the real world already.&lt;/p&gt;

&lt;p&gt;It really feels like reading a very long paper. You first want to read the title, the abstract, the table of content to gain a broad view of that the paper is about and you start with the introduction and the conclusion to make sure if you want to go into any details of the paper and then proceed with the sections of your interests. Doing early research is like do the initial steps, helping you know what is going to happen in the end and then you reading the paper chapter by chapter to build up the knowledge needed for reaching the conclusion. When you are in this process, being aware of the main objective will reduce the chances of you getting lost and help you stay focused on the important bits.&lt;/p&gt;

&lt;h3 id=&quot;how-should-i-start&quot;&gt;How should I start?&lt;/h3&gt;
&lt;p&gt;If you want to take things slowly, I think the easiest way is to take classes from different areas and go to the seminars and reading groups of that topic if you find interested. Maybe after a while you have a better idea what the field is about and might be able to identify some problems that you want to solve then you can start to approach the professor that teaches that class and then things get rolling.&lt;/p&gt;

&lt;p&gt;If you already vaguely know what interests you when entering college like my friend that I mentioned in the very beginning this post, then you can already start emailing professors to see if they have any openings for undergraduates especially the first year students. There are many things you need to know before you actually take action. The most basics include how to write an email properly. The professors aren’t like high school teachers. They are really busy. You want to try extra hard to make their life easier such that you are more likely to get a response.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;http://www.huffingtonpost.com/svetlana-dotsenko/how-to-email-professors-and-virtually-guarantee-the-response_b_9190904.html&quot;&gt;How to email professor and Virtually Guarantee the Response&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://matt.might.net/articles/how-to-email/&quot;&gt;How to send and reply to email&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://medium.com/@lportwoodstacer/how-to-email-your-professor-without-being-annoying-af-cf64ae0e4087#.im81iq8op&quot;&gt;How to Email Your Professor (without being annoying AF)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most generally, you might get frustrated if you don’t any a reply from the professor who you want to work with, send a follow up email a week after if you still haven’t heard back. Your email might just disappear in the sea of letters. The key here is really about the quality but not the quantity.&lt;/p&gt;

&lt;p&gt;When you send an email to a professor to ask for a research assistant position, you first would want to read the professor’s website thoroughly because most of the time you can find almost everything you need. Some professors put a secrete code on their websites just to filter out the students who simply spam hundreds of faculty members at the same time. There are some other professors will simply say they won’t take any undergraduate at the moment so you don’t have to waste your time sending an email.&lt;/p&gt;

&lt;p&gt;Summer is also a great time to do research, and what I found helpful was doing lab rotations. Even in the same institute you will find the group atmosphere differs and even if you already have a great supervisor already like the one I have, you will still appreciate the time you spend in other labs. Only via the lab rotation which exposes you to different things can you identify what you really like and what you don’t. It is impossible to know what the working environment is like until you are physically there being part of the group. Also, research is not a one man battle, but a collaborative effort, as a future researcher, you will for sure need to deal with people from different disciplines and countries, and the summer research experience will give you an edge over your peers for sure. Here I have a list of excellent undergraduate research programs that you can look into:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://sfp.caltech.edu/programs/surf&quot;&gt;Caltech Summer Undergraduate Research Fellowships&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://web.mit.edu/urop/&quot;&gt;MIT Undergraduate Research Opportunities Program&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.daad.de/rise/en/&quot;&gt;(German) Research Internships in Science and Engineering&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://www.amgenscholars.com&quot;&gt;Amgen Scholars&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Lastly, I want to share some excellent blogs and articles that I have been following that teach me things that I wouldn’t learn in a normal classroom setting:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;http://matt.might.net/articles/&quot;&gt;Matt Right&lt;/a&gt;: the person who wrote the Illustrated Guide to a Ph.D..&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://www.pgbovine.net/PhD-memoir.htm&quot;&gt;Philip J. Guo&lt;/a&gt;: a blog of an assistant professor of cognitive science at UCSD. His The Ph.D. Grind can give you invaluable insights.&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://da-data.blogspot.de/p/blog-page.html&quot;&gt;David Andersen&lt;/a&gt;: maintained by an assistant professor of of computer science at CMU where he shares his thoughts on various topics.&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://www.cs.unc.edu/~azuma/hitch4.html&quot;&gt;So long, and thanks for the Ph.D.!&lt;/a&gt;: a legacy guide that shall be read by all novice graduate students.&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://www.cs.cmu.edu/~harchol/gradschooltalk.pdf&quot;&gt;Applying to Ph.D. Programs in Computer Science&lt;/a&gt;: written by our beloved Mor at CMU. You can tell her passion and energy in this writing.&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://karpathy.github.io/2016/09/07/phd/&quot;&gt;A Survival Guide to a Ph.d.&lt;/a&gt;: written by Andrej Karpahthy who is a Stanford graduate student. In this article he talks about his Ph.d. experiences.&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://phdcomics.com/comics.php&quot;&gt;Ph.d. Comics&lt;/a&gt;: the must have for all!&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I didn’t know what I was supposed to do either when I was still a freshman and there were times when I thought  about why I was doing all of these. When things like this happen, make sure you have a strong support system that you can rely on. The support system can include your supervisors, graduate students in your research group, friends, family and even your college consulting center. Seek for their help. Some of them might have been in the similar situation before, so they will be able to share your sentiments and probably propose some plausible solutions and the others might have no idea of what you are experiencing, but they will at least give you mental support. Again I can’t emphasize this more, doing research or anything in life is not a one man task. You will be in much better position if you know who to ask, when to ask and how to ask help. These are the skills that you don’t learn overnight but pick up as you advance in your career.&lt;/p&gt;

&lt;h3 id=&quot;last-words&quot;&gt;Last words&lt;/h3&gt;
&lt;p&gt;The friend who pushed me to finally write this blog is Phillip Wang. We met on the first night that I arrived at CMU and I was looking for where to pick up the keys to my dorm. He was still a college freshman in the orientation week but he offered me help, seeing me looking for directions with my luggage in hand. It turned out we became good friends and I have been always impressed by his curiosity and the ability to comprehend abstract concepts that people of his age have difficulty understanding. I hope Phill will benefit from reading this and at least learn how to write a proper email instead of just using hi to address the professors : ) and of course I hope other students also have a  clearer idea of how to get started with undergraduate research.&lt;/p&gt;
</description>
        <pubDate>Sun, 25 Sep 2016 00:00:00 +0000</pubDate>
        <link>https://hangyuan.xyz/2016/09/25/undergraduate_research.html</link>
        <guid isPermaLink="true">https://hangyuan.xyz/2016/09/25/undergraduate_research.html</guid>
        
        
      </item>
    
  </channel>
</rss>
