Showing posts with label dirty data. Show all posts
Showing posts with label dirty data. Show all posts

Thursday, February 8, 2024

Flood the Dataspace


A new approach to warding off mice eating wheat seed using camouflage scents
May 2023, phys.org

Spraying wheat fields with wheat germ oil after seeding deters mice that feed on seeds.

Mice consume approximately 70 million tons of maize, rice and wheat grains each year around the globe

Colleagues in New Zealand had tried smearing the scents of endangered birds over areas where the birds would never visit. This led to predators growing suspicious of such scents, because when followed, there was no payoff. That led them to ignore the smell of the birds even when they were present.

To see if the approach might work with mice, the team treated 60 10x10 plots with wheat germ oil, which contains the scent of the wheat germ - the part of the wheat the mice want to eat. To gauge its effectiveness, the team sprayed it on plots before planting seeds and others after seeding. They also left a few plots untreated.

They were surprised to find that the oil did not serve as a false signal; the mice still ate the seeds where the plots had been pretreated. But they also found that the mice largely left alone the plots where treatment had occurred after planting. This, the researchers suggest, was likely because an overabundance of aroma had confused the mice, making it nearly impossible for them to find the seeds.

via University of Sydney: Finn C. G. Parker et al, Olfactory misinformation reduces wheat seed loss caused by rodent pests, Nature Sustainability (2023). DOI: 10.1038/s41893-023-01127-3

Post Script: In the modern world where we scavenge for data more than we do food, and in an effort to protect our own personal data, quite valuable in this modern world, we would call this "data poisoning," where you might use a different name every time you register for a webservice or purchase an e-ticket, filling the spreadsheet with similar but not the same names, confusing the predictive analytics machine with dirty data. On this note, one bonus we can expect from the deluge of artificially-generated digital detritus coming our way, is that it will fill the entire internet with fake people, which will fill the entire data-bundle of your favorite data broker with fake people, all with fake addresses and fake phone numbers and fake preferences for consumer products, and this will collapse the surveillance advertising industry.  


Friday, October 7, 2016

The Flesh of the Anthroposphere

BrainGate

The title here is in reference to the human flesh search engine, and only in theory, not in practice. The ‘flesh engine’ is used for public humiliation, but the idea of it is simply the use of a distributed process using humans as opposed to computers – not artificial intelligence exactly, but an artificially-mediated human intelligence. Nonetheless, this organizing agent, the whole of a particular population, produces a unique database of smell-descriptions that defies our contemporary notions of a “searchable database” to create the dirtiest of dirty data. 

Snippets from Hidden Scents; this one is in relation to the way People organize the names of smells, as compared to the science of chemistry or the fragrance industry:

In this compartment of olfactory classification, smell is organized by the human organism at large, the flesh of the anthroposphere, if you will. Whereas scientific taxonomies and fragrance industry inventories are themselves products of human endeavor – created for humans by humans – the (unconscious) categorization of odors by the general population is a separate system altogether.


Tuesday, September 6, 2016

When Data Has a Mind of Its Own

When good data goes bad.

Can’t hurt to post another example of dirty data, and this one is pretty cool (or not, if you’re a geneticist).

BBC News, Aug 2016

“Researchers trying to raise awareness of the issue claim that the spreadsheet software automatically converts the names of certain genes into dates.”

“Gene symbols like SEPT2 (Septin 2) were found to be altered to "September 2".”

“The researchers claimed the problem is present in "approximately one-fifth of papers" that collated data in Excel documents.”

“Excel's automatic renaming of certain genes was first cited by the scientific community back in 2004, the Baker IDI study claims. Since then the problem has "increased at an annual rate of 15%" over the past five years.

Friday, September 2, 2016

Dirty Data


I had someone ask me the other day, after reading a bit from my book, what is
“dirty data?” I was taken by surprise, because I thought that anyone under 30 knew what that was, you know, “digital natives” and all. Guess I should throw some definitions around:

Dirty data is inaccurate, incomplete or erroneous data, especially in a computer system or database. In reference to databases, this is data that contain errors. Sometimes called noise, as in signal noise, and is cleaned by a data janitor.” –wiki

Dirty data is part and parcel of Big Data and the Information Age. It’s inevitable and it’s everywhere. My autocorrect, for example, has some mis-spelled words accidentally added, and that messes up my texts, unless of course, I clean my personal dictionary. My phonebook has two different people named Nicole, obviously with two different numbers, and unless I go and disambiguate, there is no way for me to know which is which.

As a database, the “language of smell” is a stellar example of what it means to be dirty. Ask somebody what “musky” smells like, or “musty.” These are two very different smells, but because their names sound similar, people often substitute one for the other. Give someone the smell of an orange and then a lemon, and ask them which is which, but without telling them in advance. They can also both be called “Citrus.” As a database, the corpus of words we use to describe smells is a powerfully rich example of dirty data in action.

Snippets from Hidden Scents:

If knowledge is supposed to tell us which is which, and what is what, then how do we use it to study a thing that is inherently ambiguous? Smell is such a thing. In it, we have an example of an information-processing system that makes its sole purpose to ascertain ambiguous information. Moreover, during the entire process from primitive sensation to cognitive verbalization, it is fuzzy, noisy, and dirty.

Friday, July 29, 2016

So You Like Ambiguity



The duck-rabbit illusion is an oldie but goodie. This new square-cylinder illusion will make you second guess your visual cortex. It's a square, it's a circle, it's both, it's neither.


 gif:

Video:


Saturday, June 25, 2016

Age of Approximation

Screenshot from Iain McGilchrist called the Divided Brain on RSA Animate and TED.

“The Age of Enlightenment is Dead.” Thank you, Mr. Danny Hillis, for putting this in writing, and in the new MIT Journal of Design and Science no less.

Mr. Hillis makes the case, in a brief but very coherent treatise, that science is due for an update. In fact, it is not just science, but the very idea of human endeavor and progress. I recall the TED talk given by Iain McGilchrist called the Divided Brain. (This guy is author of The Master and His Emissary: The Divided Brain and the Making of the Western World, 2009). In his talk, he describes how the brain is split in two; the truth is more nuanced than that, and this is the point of the talk in fact. He goes on – there are two metaphorical sides, and they each do two different things, and the reason we tend to be right-handed is because the corresponding side of our brain is for active, intentional manipulation. And on – throughout history the pendulum swings, and one day we may regard the right side as the “right” side (because the Age of Enlightenment made us value the left side so much).

In the art classroom, I repeat these ideas, and ask them what the world would be like if everyone was “right-brained,” or artistically minded. Imagine if everyone was an artist, and nobody a scientist, all emotion and flight-of-fancy, and no bridges, tunnels, infrastructure, economic policy, institutions of higher learning, no numbers, no logic, and nothing to separate sense from nonsense. Crazy.

Then again, I can sort of imagine a world where artificial intelligence does all that stuff for us, so we can do more human things, more messy, emotional, intuitive things. I highly doubt this is what Mr. Hillis is talking about in his essay, but I’m quite excited nonetheless that he’s talking, period. He says right here:

“As our technological and institutional creations have become more complex, our relationship to them has changed. We now relate to them as we once related to nature. Instead of being masters of our creations, we have learned to bargain with them, cajoling and guiding them in the general direction of our goals. We have built our own jungle, and it has a life of its own.”

Danny Hillis, MIT Journal of Design and Science, March 2016

Post Script:
What the hell does all this have to do with the Language of Smell?

A primary objective of Hidden Scents is to present the idea that after the wave of Big Data crashes on the shores of human civilization we will have entered a new era, one in which certainty itself is no longer valued in the way it once was. We already see this today, when we ask what it is that separates us from our imminent AI overlords. Humans have intuition, something an algorithm can never have, by its nature. Humans can do this thing called “messy thinking,” or fuzzy thinking, or half-thinking. This is what leads us to make novel discoveries and connections and to be creative in general. This is what makes us not computers. And in its uncanny way, this kind of mental activity is at the core of olfaction. To smell something is to navigate a sea of data too large to fully comprehend. In this sea, one can approximate, but never ascertain. (The source of a particular smell is only verified by one of the other senses, like when you actually find that dead mouse under the fridge.)

So if the Age of the Enlightenment is dead, then perhaps olfaction (and more specifically the language of olfaction) can serve to carve the path ahead.

Well, perhaps you think I’m a bit too right-brained to be writing about such things. Thanks for reading at least.