Showing posts with label big data. Show all posts
Showing posts with label big data. Show all posts

Thursday, December 21, 2023

Coding for Information Overflow and Statistical Irregularity


Part of the "odor code" our brain uses to smell is tasked with overcoming the statistical irregularity caused by massive changes in airflow direction, speed, humidity, etc. as we pull that air through our nostrils. The cross-cancelling variables required in this effort are mentally exhausting to consider, never mind to calculate. But that's what we do when we smell:


How insects track odors by navigating microscale winds
May 2023, phys.org

"This is important because insects are typically tracking odor plumes in lower wind speeds, which indicates they are somehow making sense of the high directional variability they encounter," said Houle. "Turbulence intensity is strongly correlated with standard deviations in wind direction, which might be useful for future wind tunnel experimental designs aimed at recreating more 'natural' winds."

Based on their findings, Houle and van Breugel hypothesize an optimal range of wind speed and environmental surface complexity may exist to help insects locate an odor source.

via University of Nevada at Reno: Discovered near-surface wind direction is often highly variable over timescales of less than 10 minutes. They also found wind direction variability to be consistently higher in environments with greater surface complexity (urban areas) and lower at higher wind speeds.


Domestic cats' noses may function like highly efficient gas chromatographs
Jun 2023, phys.org

Yet another example of how in olfaction nature is still ahead of technology:

Researchers created a 3D computer model of the cat nose and simulated how an inhalation of air containing common cat food odors would flow through the coiled structures. They found that the air separates into two flow streams, where one spreads slowly above the roof of the mouth on its way to the lungs, and a separate stream containing odorant moves rapidly through a central passage directly to the olfactory region toward the back of the nasal cavity.

In essence, the researchers suggest, the cat nose functions as a highly efficient and dual-purposed gas chromatograph.

via Ohio State University: Wu Z, Jiang J, Lischka FW, McGrane SJ, Porat-Mesenco Y, Zhao K. Domestic cat nose functions as a highly efficient coiled parallel gas chromatograph, PLoS Computational Biology (2023). DOI: 10.1371/journal.pcbi.1011


Each nostril has a unique sense of smell, intracranial electroencephalogram study finds
Nov 2023, phys.org

10 subjects with intracranial depth electrodes were delivered an odor to the left, right, or both nostrils through an olfactometer device designed to deliver odors by computer control. Subjects had to identify the odor and indicate which nostril the odor came from. Subjects performed better in detecting and identifying odors in the bi-nostril condition compared to uni-nostril conditions.

Odor identity could be decoded from oscillations in the piriform cortex brain region via neural activity recorded from an intracranial electroencephalogram. The researchers observed that odor identity was encoded in two distinct, temporally segregated epochs in the bi-nostril condition, suggesting a separate smell interpretation occurs via each nostril, suggesting a possible computational advantage in processing odors in stereo. 

via University of Pennsylvania and the Barrow Neurological Institute of Phoenix: Gülce Nazlı Dikeçligil et al, Odor representations from the two nostrils are temporally segregated in human piriform cortex, Current Biology (2023). DOI: 10.1016/j.cub.2023.10.021

Thursday, January 16, 2020

Sensory Nutrition




The Monell Center for taste and smell research sheds some light on the emerging field of Sensory Nutrition. Sure we're all human, and all made of the same stuff, and all programmed by DNA that is pretty darn similar. But we are not the same. We don't even taste or smell things the same, and much of that difference starts with our DNA.

Boy did I have a great conversation the other night about a friend of a friend who tried to fuse Mexican food into a Korean city's cuisine. Didn't work. Why? Cilantro, that's why.

Ambitious food alchemist didn't do his homework -- Asians in general tend to taste cilantro as "soap," i.e., gross. This isn't about preference, it's about genetics. For whatever reason, some of us code cilantro as soap and others as the most refreshing herb ever.

Monell researchers could have told him that. They're using big data, machine learning, and genome-wide association studies (GWAS) to understand the interface between sensory science, nutrition, and dietetics. They're ultimately trying to see if we can guide people into the right public health intervention just based on their genes.

Behavioral geneticist Danielle Reed, PhD, and olfactory neurobiologist Joel Mainland, PhD, helped to mine 400,000 reviews of 67,000 food products posted by 256,000 Amazon customers over 10 years. That's the big data part. The machine learning part analyzed words related to taste and smell, as well as other categories related to health.

Output? People today think food is too sweet. Wow. Never would have guessed that. No matter kind of food they were talking about, one percent of all reviews used the words "too sweet."

On the other side of the taste spectrum, and from a totally different study – there's a gene that helps you taste bitter, but if you have a hyped-up version, you will taste too much bitter, especially in vegetables like dark greens. Maybe even other bitter things coffee and beer will taste way different to you.

For reference (go ahead, dial up your time machine to about ten years into the future and pull up your DNA database), it's the taste gene TAS2R38. It codes for bitter-taste receptors on the tongue. And it has two variants, the AVI and PAV variants. Depending on the combination, you'll have a very different experience with certain bitter chemicals.

So the headline is that we're hardwired to like or dislike vegetables. Camouflaging bitter tastes with culinary creativity might not hurt. Just make sure to do your homework.

Notes:
Monell Center, Philadelphia PA

Nov 2019, BBC News

Monday, August 14, 2017

Place Your Bets




Mr. Avery Gilbert, sensory psychologist, author of lots of books on smell, and writer for the blog First Nerve, asks the hard questions about smells. For those who aren’t on the science side of smell, this topic might be too much. But really, it’s getting to the bottom of how smell works, and how we might be able to categorize this sense that seems to be impervious to organization. Spoiler alert – it doesn’t work, and we’ll try to explain why.

Gilbert asks, “Can we predict a molecule’s smell from its physical characteristics?” This is an important question that has lots of people wishing it were true. The problem is that we can’t seem to predict what something would smell like, given any information about its molecular properties. Sure, sulfur-things smell sulfur-y. But there is a lot more to smells than that.

There is a dream out there, that a periodic table of smells really exists, that there is a chart listing all the molecules and all their smells, and all are organized into neat rows and columns. But that’s just impossible. There are infinite molecules and infinite smells. Further, there are infinite ways to describe these already-infinite things because the language of smell, as it tries to pin-down and describe, only begets more words, not less. So at the outset, this is impossible.

But still, Science is trying, and recently (this article is from February), a huge crowdsourced effort claims to have found a pretty good means of doing this. It’s called the Dream Challenge, and it uses a bunch of parameters, or features of molecules all thrown together in a big data pile, in the hopes of returning an algorithm that can smell as good as we can, but just by looking at the molecular data. These features, “4884 physicochemical total features of 338 molecules,“ might be the chemical constituents of the molecule: hydrogen, carbon, etc.; or the family of chemicals: ketones, aldehydes; or something about their three-dimensional arrangement. All these parameters are put together, and a big computer sifts through them looking for a pattern between these features and the names people give them, in regards to their smell.

So this brings us to the next part, which is the more important part, for me at least.  What do we call these molecules and their perceptive impressions? In other words, where do the names of the perceived smells come from? Scientists for decades have been using a database of smell names and their matching molecules that comes from the 1980’s and was put together by a man named Andrew Dravnieks. The odor atlas it’s called, and it’s been used for every experiment dealing with smell. It’s pretty much the only one. For the Dream Challenge, a new odor atlas was created. And therein lies the rub, for Avery Gilbert, as well as for anyone who is the least bit suspicious of science predicated on a subjective lexicon.

The new psychophysical dataset was created like this – they gather 50 people to be the smeller/namers. This is a good number, as Mr. Gilbert says, because it cancels the perceptive variability that is bound to occur between people (because for many reasons from genetics to Freudian psychoanalytics, no two people smell the same thing). Then these people are given a bunch of odorants, molecules, and asked to describe what they smell like.

These smellers don’t get to name the odorant molecules whatever they want; they are given a list of pre-determined words, like fruity or burnt. This is the first catch. In this case, the predetermined words listed only 19 (garlic, sweet, fruit, spices, bakery, grass, flower, sour, fish, musky, wood, warm, cold, acid, decayed, urinous, sweaty, burnt, and chemical). That’s more than some other studies have tried to use, but less than the available lexicon, and still less than some other recognized lexicons, like the ones used in wine (86 descriptors) and coffee (85) and the Dravnieks odor atlas (146).* Then there’s the descriptive capacity of the words chosen. “Urinous” is pretty specific, but “cold” is not. (What the heck does “cold” smell like anyway? The opposite of “warm, ” of course!)

The reason the number of descriptors is set at 19 is not because there are only 19 ways to describe the hundreds of odorants offered. The reason is because the set of all possible descriptors has been collapsed, or organized into categories. When creating a lexicon like this, many descriptors would tend to be similar – they are near each other in odor-name-space. But this can be an illusion that is really hard to penetrate. ‘Musty’ and ‘urinous’ seem like they might go together, no? Or ‘sweet’ and ‘fruity?’ And how many other things fall under the ‘garlic’ category to actually constitute it as a category? Can you call something a category if it only has one thing in it?

Avery Gilbert explains that many of these lexicons are made for specific purposes, imposing a structure on an otherwise cacophonic mess. The wine wheel is for tasting wine, the coffee dataset for tasting coffee, etc. In this case, there was no purpose; it’s supposed to be universal. And with this total lack of context, the semantic descriptors available may not directly correlate to what the smeller really smells.

Bottom line is, if this Dream Challenge comes up with an algorithm to predict what something smells like based on its chemical properties, it will be surprising.

It’s just one of those things that seems like it can’t be done. It’s one of those things that is out of its element, like a gorilla trying to mate with an elephant and expecting a healthy offspring. A psychophysical dataset is this, a gorillaphant. Chemistry is strict, rigid, objective, distinct, and quantifiable. But words are polysemous, they mean more than one thing. And smelling as a perceptive act is too personal. There is no standardization to the way we name smells. There is no Pantone for smells. We never ever learn the names of smells in a standardized way, so how can we have a universal dataset for them?

Personally, I think the National Geographic Smell Survey is a more interesting study. This is something Avery Gilbert worked on. Because it covers people all over the world, it gets a better sense of what a smell is. Smells are hard to name because they are subjective, for one, but also because the names we give them are very culturally determined. To survey people all over the world (as in the NatGeo survey) is a good way to cross cancel the problem. Then there’s the time issue. How long does something like that last? This survey was done in 1989. That’s around when Calvin Klein came out with “unisex” perfume, and way before they sold kimchee in my non-asian neighborhood supermarket. Culture changes, both across time and space, and the words we use to describe our experiences change with it. We may never have a universal smell chart, not until we ourselves are universal. (And hopefully that never happens.)


* U.C. Davis Wine Aroma Wheel uses 86 terms to describe wine, the World Coffee Research Sensory Lexicon uses 110 terms to describe coffee and the Dravnieks odor atlas uses 146.


Journal articles referenced:

“Predicting human olfactory perception from chemical features of odor molecules,” by Andreas Keller, et al., published online February 20, 2017 in Science.

“Olfactory perception of chemically diverse molecules,” by Andreas Keller and Leslie B. Vosshall, BMC Neuroscience 17:55, 2016.

National Geographic Smell Survey. Wysocki C, Gilbert AN. Ann N Y Acad Sci. 1989;561:12-28.



Image source: link

Monday, August 7, 2017

Brits and Twits


Some words just sound funny, and sometimes sounds are just more important than semantics.

Booty, booby and nitwit—academics reveal funniest words
Aug 2017, phys.org

Tomas Engelthaler and Professor Thomas Hills in the Department of Psychology analysed 5000 randomly selected words, and showed them to over 800 people online - asking them to rate from one to five how humourous they found each word.

The funniest words - those which had been scored highly by the most people, and given the highest mean humour rating - were (in order):

Booty
Tit
Booby
Hooter
Nitwit
Twit
Waddle
Tinkle
Bebop
Egghead
Ass
Twerp

-phys.org

I have a feeling the sample pool here is from England not the US. Something about 'twit'.

Post Script
Sound Symbolism and Universal Language
Limbic Signal 2017

Sunday, July 23, 2017

Constructing an Urban Smell Map



Must be something in the air. Over the past few years, many artists/scientists have been making smell maps of their cities. Victoria Henshaw is the most popular, with her “smellwalks” and her book Urban Smellscapes. But there’s others, like Kate McLean’s Sensory Maps, and Jason Logan’s handscrawled Scents and the City (seen above).

I think it’s just that maps are becoming a big deal. Something about big data, social media, and GIS interface. Maybe that Snapchat debacle too. The smell map we’re looking at today wraps up all three of these, and makes a very ambitious project into a neat interactive plaything. I wasn’t satisfied with the map, however, and was way more interested in the Urban Smell Dictionary they created in order to make their map I the first place. I was suspicious, so I went through their full report.

So here goes:

First they conducted smellwalks in each of the cities to be mapped (London and Barcelona), and along with some previous literature on the subject, they used the words generated on these smellwalks to support their lexicon. Then they collected geo-referenced picture tags from Flickr (530K), Instagram (35K), and Twitter (113K). Those tags and tweets were matched with the words in the smell dictionary, and voila!

The lexicon, in detail:

First they did some co-occurrence modeling, so that smell-words that appear in the same post/tweet are related. Then they used network analysis software, also called a community structure detection algorithm. InfoMap is the one they used, but they did further partitioning with another program.

The result?

Urban Smellscape Aroma Wheel



Who did it?
University of Turin computer science professor Rossano Schifanella, and Bell Labs researchers Luca Maria Aiello and Daniele Quercia.

Image source: “I Smell NY” Smell Map by Jason Logan

Notes:
Quercia, Daniele, Rossano Schifanella, Luca Maria Aiello, and Kate McLean. 2015. “Smelly Maps: The Digital Life of Urban Smellscapes.” In Proceedings of the Ninth International AAAI Conference on Web and Social Media (ICWSM 2015). Palo Alto, CA: AAAI Publications. 327-336. Accessed November 28, 2016.

Full pdf:

Post Script:
There was a really cool observation made during their report which I’d like to explain here. In the remarks they make about “base notes,” there's some good entropy data in there somewhere, about smells, awareness, and the amount of time you spend in a place/some measure of odor-attenuation.

Let's start with their explanation of "notes" (and I’ll just quote from the paper)

“When a new perfume is created, different top, middle, and base note ingredients are combined to make the new fragrance. Those notes differ in terms of their tenacity. Top notes are those perceived immediately (e.g., citrus fruits, aromatic herbs) and, since they are intense, they are also volatile and evaporate quickly. By contrast, base notes are those adding depth and stay on the skin for hours (e.g., wood, moss, amber, and vanilla). Middle notes sit somewhere in between (e.g., flowers, spices, berries).”

Now they explain the urban smellscape:

“Base notes. The macro-level base notes for the urban smellscape are those that are likely smelled by a city’s first-timer visitors. That is because known odors are unconsciously processed by people, while only unfamiliar or strong odors are brought to people’s attention (as potential threats or sources of pleasure). As a result, residents are not likely to pay attention to their city’s base notes, while visitors would be able to consciously process them.

“Mid-level notes. As one moves through the city, the base notes blend with dominant smells that are localized in specific areas (e.g., factories, fish markets).

“High notes. Finally, the micro-level high notes are shortlived odors (e.g., goods from a leather shop). These are emitted in points that are very localized in space and time. … High notes are likely to go undetected because of data sparsity and because of our spatial unit of analysis being a street segment.”

And one more thing, I’m like how do you say this and not cite it??

 “Air pollutants also have been found to reduce the ability of floral scent trails to travel through air.”

-I’ve never heard this, and they provide no citation

Tuesday, July 18, 2017

Corpus of Hedonics



Looking at this paper today:
The Emotional and Chromatic Layers of Urban Smells. Daniele Quercia, Luca Maria Aiello, Rossano Schifanella. 2016.

The chart above shows the relation between a smell-producing location, and the pleasantness of the words used on social media near that place. Most positive words are used near the Food category, and less near Waste. Nature has less happy words with it than Food, perhaps because of the happy activities people are doing in these areas, i.e., eating and being with friends. Regardless, this chart seems to make sense, and can serve as proof that this odor-hedonics data can be predictive of people's emotions.

Who cares? Someone like an urban planner can use this to get an idea of how to better arrange municipal facilities. Someone who wants to update local land use regulations could use this data to see which areas of town work and which ones don't. The study-authors mention the matching of this data to 'most optimal route' data to give the 'most pleasurable route' through a city.

They used a lot of semantic-hedonic smell knowledge from a previous study (Henshaw 2013) which organizes two lists of pleasant and unpleasant smells, and I'd like to just copy it here:

Pleasant smells
bread, baked, baked goods, coffee, coffees, aftershave, cut grass, grass, grassy, floral, flower, flowers, flowershop, flowery, lavender, lilies, lily, magnolia, rose, rosey, tulip, tulips, violet, violets, baby, babies, child, children, sea, seaside, countryside, cedar, cedarwood, conifer, dry grass, earth, earthy, eucalyptus, ground, leafy, leaves, old wood, pine, sandalwood, soil, tree, trees, wood, woodlands, woody, petrol, diesel, fuel, gasoline, soap powder, soap

Unpleasant smells
flatulence, fart, vomit, dog shit, dogshit, excrement, faeces, farts, feces, manure, shit, cigarette smoke, cigarette, cigarettes, cigar, cigars, smoker, tabacco, tobacco, pee, piss, ammonia, urine, public toilet, public toilets, toilet, toilets, urinal, urinals, gone-off milk, fish, rotten fish, rotten food, rotten, rotten fruit, rotten fruits, putrid, bus, buses, car, cars, exhaust, traffic, fume, fumes, body odour, body odor, sweat, sweaty, dirty clothes
  Henshaw, V. 2013. Urban Smellscapes: Understanding and Designing City Smell Environments. Routledge.

Post Script:

This is how they got to the bottom of their smellscape emotion chart:

"We set out to study the relationship between the smellscape and emotions on our data. To do so, we need to have a lexicon of emotion words. We use two of them: the “Linguistic Inquiry Word Count” (LIWC) (Pennebaker 2013), that classifies words into positive and negative emotions, and the “EmoLex” word-emotion lexicon (Mohammad and Turney 2013), that classifies words into eight primary emotions based on Plutchik’s psycho-evolutionary theory (Plutchik 1991) (i.e., anger, fear, anticipation, trust, surprise, sadness, joy, and disgust).
  Pennebaker, J. 2013. The Secret Life of Pronouns: What Our Words Say About Us. Bloomsbury.
  Plutchik, R. 1991. The emotions. University Press of America.
  Mohammad, S. M., and Turney, P. D. 2013. Crowdsourcing a word–emotion association lexicon. Computational Intelligence 29(3):436–465.

Wednesday, February 22, 2017

Non Traditional Computing, Complex Problems, and Approximation

Illustration for Death of a Salesman by Brian Stauffer for the Soulpepper Theater Company, Toronto

It may seem like a stretch to write about combinatorial optimization problems (aka the traveling salesman problem) on a blog about the ‘language of smell,’ but Limbic Signal isn’t just about smells, or language, but the connections between olfaction and computation. Our olfactory system is a champion at dealing with very large, very complex datasets.

Olfaction uses our brain in ways the other senses don’t. Some of the ways olfaction diverges from the other senses are akin to novel solutions to very complex problems in computation, such as big-data-sifting, pattern recognition, or the aforementioned traveling salesman problem.

 Also note that, in addition to the magnet network described below, another unconventional solution to the traveling salesman problem is to use mold. In fact, slime mold was used to design Spain's motorways and the Tokyo rail system.

So this article below does a good job of explaining the traveling salesman problem; I straight copied it from the writers at phys.org. And in the second section is an explanation of an interesting solution to the problem.

Researchers create a new type of computer that can solve problems that are a challenge for traditional computers

The traveling salesman problem
There is a special type of problem - called a combinatorial optimization problem - that traditional computers find difficult to solve, even approximately. An example is what's known as the "traveling salesman" problem, wherein a salesman has to visit a specific set of cities, each only once, and return to the first city, and the salesman wants to take the most efficient route possible. This problem may seem simple but the number of possible routes increases extremely rapidly as cities are added, and this underlies why the problem is difficult to solve.

...
It may be tempting to simply give up on the traveling salesman, but solving such hard optimization problems could have enormous impact in a wide range of areas. Examples include finding the optimal path for delivery trucks, minimizing interference in wireless networks, and determining how proteins fold. Even small improvements in some of these areas could result in massive monetary savings, which is why some scientists have spent their careers creating algorithms that produce very good approximate solutions to this type of problem.

An Ising machine
The Stanford team has built what's called an Ising machine, named for a mathematical model of magnetism. The machine acts like a reprogrammable network of artificial magnets where each magnet only points up or down and, like a real magnetic system, it is expected to tend toward operating at low energy.

The theory is that, if the connections among a network of magnets can be programmed to represent the problem at hand, once they settle on the optimal, low-energy directions they should face, the solution can be derived from their final state. In the case of the traveling salesman, each artificial magnet in the Ising machine represents the position of a city in a particular path.



Tuesday, September 6, 2016

When Data Has a Mind of Its Own

When good data goes bad.

Can’t hurt to post another example of dirty data, and this one is pretty cool (or not, if you’re a geneticist).

BBC News, Aug 2016

“Researchers trying to raise awareness of the issue claim that the spreadsheet software automatically converts the names of certain genes into dates.”

“Gene symbols like SEPT2 (Septin 2) were found to be altered to "September 2".”

“The researchers claimed the problem is present in "approximately one-fifth of papers" that collated data in Excel documents.”

“Excel's automatic renaming of certain genes was first cited by the scientific community back in 2004, the Baker IDI study claims. Since then the problem has "increased at an annual rate of 15%" over the past five years.