Showing posts with label tools. Show all posts
Showing posts with label tools. Show all posts

Wednesday, January 11, 2023

Better article PDF

When you publish an article, you want it to be discoverable by other researchers. This requires that it be indexed. Indexing systems need metadata, which they usually extract from the PDF files in digital libraries or other document repositories.

Ideally, a manuscript submission system collects the necessary metadata from the corresponding author when the final manuscript is submitted. Since author names are not unique, a good system will require each author to log into the system using their ORCID credentials, which also ensures that all coauthors know they are such.

Some metadata is not known by the submitter, for example, the DOI, volume and issue number, page, and publication date. Such metadata is added by the managing editor. In some journals, the managing editor creates all metadata, but it is safer when the publication system generates algorithmically the metadata from that provided by the corresponding author.

Some publishers omit the verified ORCID collection from the authors. In that case, you should put each author's identifier in the author declaration. When you create your manuscript using LaTeX, you can use the package orcidlink. In the preamble add

\usepackage{orcidlink}

and in the author declaration add your ORCID

\author{John Doe\,\orcidlink{nnnn-nnnn-nnnn-nnnn}}

Some publishers just publish in their digital libraries the final PDF they receive. If you just submit the default format file you generated, it will not have any metadata and your article will not be discoverable. It is safer, to submit the article in the PDF/A archival format and with your metadata included as XMP so it can be extracted by the web crawlers of the indexing organizations.

You could accomplish this using the full version of Acrobat, but as I wrote above, manual operations are not recommended and you should let the LaTeX typesetter do it. Fortunately, River Valley Technologies has contributed a package called pdfx to automate this step. You already have this package with the standard LaTeX installation.

After importing this package, you also import hyperref. Then you write an xmpdata file declaring the metadata. You can find all the information in the exhaustive help file. The import order is important, for example in the preamble you could declare

\usepackage[a-1b]{pdfx}

\usepackage{hyperref}

If you are an editor and maintain a template for your journal, you can also embed the xmpdata file at the top of the preamble. The help file explains how to do that.

You can check the PDF format and the metadata with the free Acrobat Reader. I have generated the following two documents using the method described above:

article

slides



Monday, June 21, 2021

Industrial Research

Previous related posts: Career Networking, Scholarly Publications.

With the industrial revolution, large companies introduced central research laboratories to accelerate the invention of new products. After the Sputnik crisis, in the USA these laboratories became very prestigious, under the influence of people like Joseph Licklider. The labs flourished, receiving government jobs paid under the cost-plus model, where the paid price was the cost of producing the technology plus a margin for the company's profit.

After receiving their Ph.D., young researchers would join learned societies in their field and attend their annual meetings to stay current in the field and to network. They would join a prestigious central lab with the plan to work there until retirement. With the pressure for excellence, there were always personal frictions, but they were offset by the companionship and respect researchers had for each other.

In research, the hierarchies are relatively flat. The organizations were very flexible, and researchers moved around in the lab and formed robust networks. They would attend the annual meetings of their societies and present their progress: their network was not only dense but also vast. By subscribing to the same journals, there was a shared knowledge of the state of the art.

At the end of the Cold War, the central labs rapidly disappeared. The government no longer had the need to outbrain the Russians, companies put more emphasis on quarterly results instead of long term success, therefore executives had less understanding for research. The universities adapted and started programs to teach students to become entrepreneurs. All this contributed to the central labs to disappear in a very short time.

At first, one would think that industrial research has completely disappeared. However, this is not the case, and today there are more researchers than during the Cold War. But the research infrastructure has radically changed. For example, professors no longer spend occasional time in industrial central labs as visiting scientists, but have more secure part-time positions in large companies, with job titles like Fellow.

Typically, a larger technology company has a VP of research with numerous researchers. The latter no longer sit in a central location but are dispersed throughout the engineering divisions. On one side, this allows them to glean important hard problems with which the engineers are grappling and get inspired for new technologies. On the other side, when engineers get stuck with a problem for which there is no clear solution on Stack Overflow, they can informally ask the local researcher for a lead.

The researchers are not embedded in the development organizations. They report to a remote manager, and they have long term goals instead of the daily Jira tasks. They do not have short term deadlines, but at the end of the day they cannot turn off their brains until the next morning. Today's researchers are much more lonely than the researchers in the central labs of yore.

Today's researchers are less dependent on learned societies and tend to network using LinkedIn. The annual meetings have disappeared and have been replaced by topical meetings. Instead of a steady shower of scientific articles in journals, researchers today do searches on Google Scholar for the knowledge they need at the moment. One of the corollaries is that today papers should no longer have memorable titles, but the titles have to have the important terms at the beginning, so they show up at the top in searches.

This requires learned societies to adapt. Researchers are isolated and change employers more often. Flexible regular meetups are maybe more important than rigid conferences. When in the past societies could solicit sponsorships from central labs, today the researchers are decentralized and there is no longer a budged item for sponsorships. Despite this, anecdotally there is more money for essential expenses and while in the past page charges were a barrier, today they are no longer important and researchers publish in journals with high impact factor, regardless of cost.

There is another important factor in the lives of researchers. Since they are now dispersed, it is more difficult to advance in the career. The Anglo-Saxon countries always had the concept of mentorship instead of the more formal master-apprentice system of other western countries. Other cultures are now copying the mentorship system to help researchers to succeed in life. This is a new role to which learned societies have to pay attention.

As I mentioned, researchers do searches for related art on the web instead of staying current by subscribing to journals. Search engines are based on n-grams and do not know the history of science. Therefore, it is easy to get the related art in the introduction wrong, and especially the references are often wrong. For example, the CIELAB color model operator was not introduced at the INTERACT-2010 conference, but much earlier and no later than 1976.

Thus, editors and reviewers have a much harder job verifying the introduction and the references of a manuscript. This and the current career path of researchers (see post on Career Networking) prompted, for example, the SPIE about 15 years ago to limit the terms of the editors in their journals. Learned societies must give high consideration to mentorship for their journals and conferences. Conference chairs and editors must have a good number of young researchers who work closely with the old hands to learn the ropes. Fortunately this is easy because, as noted earlier, today's researchers are lonely and long for mentors.



Sunday, June 20, 2021

Scholarly Publications

The evolution of approaches to employment was discussed here.

Fifty years ago, a professor in the mathematics department at the Swiss Federal Institute of Technology (ETH) was expected to publish a substantial paper every two years, at least one every 4 years. The paper was top quality and presented a true advance in the field. Once a year, the professor would also present at a conference. When the professor had come up with a lecture presenting a new field or a novel approach to an existing field, their Ph.D. students would sit in the front row of the auditorium and take extensive notes, which would be the basis for a new book.

Students were not expected to write papers. They would write reports and give seminar presentations to get the required credits.

The experimental physics departments was quite different: the American publish-or-perish way of life had taken over. By twenty years ago, also the mathematicians were living by the publish-or-perish paradigm. However, something else changed: the students would submit their reports to scholarly journals.

The flood of submissions required the editorial boards to change their criteria for reviewing manuscripts. Moreover, with the introduction thirty years ago of the World Wide Web by the European Organization for Nuclear Research (CERN) in Geneva, a huge number of publications became easily available, making the editor's job of separating the wheat from the chaff more urgent, otherwise researchers just waste their time reading useless articles.

Certain measures were easy to implement, like eliminating unintelligible manuscripts, plagiarisms, and nothing-wrong-papers (papers well written but not advancing the field). More editorial work was required for papers where the authors did some valid research, but did not understand it well themselves—in principle, author's supervisor would be responsible, but for the past twenty years they have been increasingly remiss of this duty. The other editorial task is to identify salame papers and reject them; salame papers refers to when the result of a project is sliced up and submitted as a series of papers. While this is OK for conference papers, it is not for journal papers.

A form of plagiarism was also to submit a paper to multiple journals using variations of the author's names. This was solved by requiring a persistent digital identifier (ORCID iD) and using a Digital Object Identifier (DOI) for every reference.

When fifty years ago papers were well written, today they tend to be sloppy. When an editor accepts a manuscript for publication, the authors tend to ignore the orthography and grammar errors pointed out by the reviewers, even when good authoring tools are available. This sloppiness increases the publication cost because a copy editor has to rework the manuscript. When a sweatshop is used, producing a 12-page article typically costs about $1,500 while using professional copy editors doubles or triples the cost. Societies usually slightly increase the publishing fee to allow for discounts for members and to subsidize financially challenged authors.

Fifty years ago, a scholarly journal covered its production costs by charging a small page charge and with subscriptions by institutional libraries. With the flood of articles in the past twenty years, the number of journals has increased and with the higher required production costs, libraries have an issue affording journals. In the long term, scholarly journals are only viable with the open access approach, where the authors pay the full publication costs.

For a large institution, the publication costs are significant especially when the copy editing cost go up with the increasing sloppiness. This leads to a hybrid solution where institutions pay a fixed yearly price and get a certain number of submissions and downloads.

Usually institutions are not too sensitive to the publication costs, as long as the journal has a good impact factor. A journal builds its impact factor not with the work of the copy editors, but with the work of the editors. For scholarly journals, these are usually researchers who volunteer their work. The question is how does one find a good team of editors-in-chief and associate editors? For this we have to look at how research evolved in the last decades.

Before, let me point out that editors have to ensure that relevant references to articles in their own journal should not be omitted and the journal must be indexed. Last but not least, the most important words of the articles should be at the beginning of the title, otherwise citing authors will miss it in a Google Scholar search.



Sunday, June 2, 2019

Efficiency

One of the main avenues to increase the quality of life and the welfare of society is to become more efficient. Concerning one of our most labor-saving devices, the computer, in the last decade, we have not done so well. While previously it had liberated us from typewriters and slide rulers, recently it has been distracting us with social media and CPU power has barely improved, although computers have become pocketable and are now ubiquitous.

However, it is still worthwhile to periodically update a workstation, for example by using a more powerful graphics card with a better GPGPU. Another worthwhile upgrade is to replace hard disk drives (HDD) with solid state drives (SSD). These have become very economical and reliable. From a hardware point of view, the easiest upgrade is to replace the internal 3.5" SATA disks with 2.5" SATA SSDs: it just takes a couple of minutes.

Open the workstation, slide out the drive tray (left), remove the disk, screw in a 2.5" to 3.5" SATA adapter (right), screw the SSD (middle) into the adapter, then slide the tray back in the workstation and close the lid.

replacing a hard disk drive with a solid state drive

The software side is a little more complicated and takes hours (but you do not have to watch it). Modern operating systems have a special partition on the system disk called GUID for the firmware. On some operating systems this partition contains the drivers, including the file system code. On my older workstation, the firmware is in a PROM, but the GUID partition is still required because the EFI system partition is used as a staging area for firmware updates. See this article for the macOS.

The safest update procedure for the software is to format the new drive and then do a fresh install of the operating system. This will start with a firmware update (press the power button until the power light flashes and you hear a long beep). In my case, the firmware was updated from version MP51.0087.B00 to version MP51.0089.B00. If you do not update your boot ROM, from time to time you will get DiskManagement error -69546 from macOS.

With the new firmware, reboot the workstation and install the operating system. In the end, do a full system migration from your backup disk. Your system is now much faster, especially when you boot it up. While I was at it, I also replaced my backup disk with a very inexpensive but high-quality consumer-grade SSD shown below.

an inexpensive consumer grade SSD can be used for backup

No other changes were required and the workstation has now become much more efficient.

One aspect that has become very inefficient in the last years is buying parts. We used to be able to go to the neighborhood store and find everything, but nowadays these stores are in bad shape and it has become difficult to find items due to online stores. In the case of the SSDs, it was not an issue because instead of picking them up in the store in 20 minutes I got them in a couple of days from the manufacturer.

However, for the 2.5" to 3.5" SATA adapter, I was less lucky. The electronic store had over a dozen different adapters, but there was no way to find out which one was for my workstation. I figured that if I order online I will get it in a couple of days from a warehouse in the Central Valley or in Utah, but I was in for a surprise. It came from overseas and unfortunately for the shipping company Palo Alto is near New York and the adapter went on a long random walk across the continent. I feel really bad for my huge carbon footprint to ship this small $16 part.

date / time activity
Wednesday, April 24, 2019 11:42 AM Received electronic information
Friday, April 26, 2019 10:57 PM Shipment information sent To FedEx
Saturday, April 27, 2019 10:49 AM [China-Shanghai Operations Center] waiting for transshipment
Saturday, April 27, 2019 11:47 AM [China-Shanghai Transit Center] left scanning
Sunday, April 28, 2019 4:58 AM [China-Shanghai Transit Center] left scanning - loaded car
Sunday, April 28, 2019 7:06 AM [China-Shanghai Pudong International Airport] Arrival at the airport - exi
Sunday, April 28, 2019 8:06 AM [China-Shanghai Pudong International Airport] Customs Release - Export
Sunday, April 28, 2019 12:17 PM [China-Shanghai Pudong International Airport] parcels from developing countries
Monday, April 29, 2019 10:03 AM [United States - Kennedy Airport] arriving at the airport - import
Thursday, May 2, 2019 11:41 AM [United States - Kennedy Airport] Customs Release - Import
Friday, May 3, 2019 7:47 PM [FEDEX SMARTPOST BREINIGSVILLE, PA]Arrived at FedEx location
Saturday, May 4, 2019 5:50 AM [FEDEX SMARTPOST BREINIGSVILLE, PA]Departed FedEx location
Saturday, May 4, 2019 10:34 PM [JEWETT, IL]In transit
Sunday, May 5, 2019 10:41 AM [QUAPAW, OK]In transit
Sunday, May 5, 2019 9:44 PM [SAN JON, NM]In transit
Monday, May 6, 2019 8:54 AM [TOPOCK, AZ]In transit
Tuesday, May 7, 2019 2:04 AM [WALNUT, CA]In transit
Tuesday, May 7, 2019 2:11 PM [BAKERSFIELD, CA]In transit
Wednesday, May 8, 2019 9:53 PM [FEDEX SMARTPOST SACRAMENTO, CA]Arrived at FedEx location
Thursday, May 9, 2019 12:27 AM [FEDEX SMARTPOST SACRAMENTO, CA]Departed FedEx location
Thursday, May 9, 2019 2:42 AM Shipment information sent To US Postal Service
Saturday, May 11, 2019 Delivered

Monday, October 22, 2018

Career Networking

In the US, the multigenerational workforce is divided into five age groups, which have quite different approaches to employment.

The traditionalists (or silent generation, born 1925–1945), have these stereotypical characteristics: striving for financial security; "waste not, want not"; nobility of sacrifice for the common good; focus on quality and simplicity; loyal to employers and expect loyalty in return; believe promotions, raises and recognition should come from job tenure; work ethic focused on timeliness and productivity; conformity and following authority.

The baby boomers (born 1946–1964), have these stereotypical characteristics: the importance of hard-work (instilled by parents); loyalty to an employer would lead to reward and seniority; willingness to take on additional responsibilities; conscientious and dependable; service-oriented; ambitious; dutiful.

The generation X (born 1965–1981) have these stereotypical characteristics: the importance of education; shaping one's own career path; work-life balance and autonomy; innovation and entrepreneurialism; comfortable with challenging conventional wisdom; outcome-oriented; collaborative decision making.

The millennials (or gen Y, born 1982–1997) have these stereotypical characteristics: need intellectual challenge; entrepreneurial; value continuous learning opportunities; achievement / results-oriented; innovative and open to new ideas; collaborative decision makers; like praise and recognition; value teamwork and equality; value independence / autonomy; seek meaningful work; value work-life balance and flexibility; value fun at work; technology-driven

The centennials (or iGen or gen Z, born 1998 and later) are just entering the workforce and the stereotypes have not yet been formed.

The traditionalists rely on local organizations like the Rotary or the golf club for networking, but also professional societies and conference attendance. The baby boomers participate actively in international societies and conferences, building a global network. The generation X still participates in conferences but is less active in professional societies and the organization of conferences. The millennials are on social media and use search engines to find information and attend local meet-ups for networking.

While in the past peoples managed contacts using a Rolodex, membership directories, etc., today colleagues are constantly on the move and everybody has to maintain their personal contact information on a professional network like Xing or LinkedIn, through which they connect to their professional contacts.

Professional network sites make money by selling your information to business intelligence and salespeople as well as recruiters looking for employees. The service is free for you, but you have to maintain your own information.

The sites are continuously improved, so you have to keep monitoring your profile for changes in the way your information is organized to be more valuable to paying customers. For example, the skills section is sorted by the number of endorsements you receive for each skill, which is not what you want. Edit this section by clicking on the pencil on the top right, then unpin the top three skills, reorder the skills by dragging the horizontal lines on the right, and pin your top three skills.

When LinkedIn bought SlideShare, your presentations appeared in the media section. However, the original site was abandoned and your media is in cold storage. To get acceptable access times, you have to upload your PDFs again directly into LinkedIn. Furthermore, videos are no longer supported, so you have to upload them to YouTube and then make them available in your LinkedIn profile as linked media.

If you apply to a job on a professional social network by clicking on the "apply" button, the probability that you will have that job in your profile is very low. Instead, you have to click on your best connection working there because jobs go mostly through internal referral.

Of course, you have to have a contact working there. The quality of the contact is important because this person has to be your advocate. You can easily increase your network by turning on Bluetooth in your mobile LinkedIn app and invite all people in your vicinity, but they will not be your advocates. Your network has to be dense.

LinkedIn connection map

The best way to create a dense network is to organize conferences because people will remember well your skills and leadership qualities. The second best is presenting at conferences and the easiest is to present at meet-ups. Even easier is to write a blog, but you should post at least once a week and advertise each post on LinkedIn and Twitter.

Friday, January 12, 2018

Annotating detected outliers

The so-called Twitter Anomaly Detection function for R is excellent but also very minimalistic. The input is a two-column data frame where the first column consists of the timestamps and the second column contains the observations. In addition to a plot, the output is a data frame comprising timestamps, values, and optionally, expected values.

In practice, we usually have some semantic information that we would also like to include in the output, so we do not have to refer back to the original data. Fortunately, there is a quick-and-dirty way to add a description to the outlier data frame.

We start with the annotated data frame containing at least columns with the timestamps, the observations, and factors providing contextual or semantic information on each observation. We then create a simple data frame with just the first two columns, which we pass to the outlier detection function.

We can write a trivial function that for each outlier finds the row index in the simple data frame and looks up the semantic information in the annotated data frame:

AddDescription <- function(series1, series2, outliers) {
 quantity <-  lengths(outliers$anoms[1])
 if (quantity < 1) return (NULL)
 else {
   result <- NULL
  for (i in 1:quantity) {
   rowIndex <- which(series1$timestamp == outliers$anoms$timestamp[i])
   newRow <- data.frame(outliers$anoms$timestamp[i],
    outliers$anoms$anoms[i],
    as.character(series2$note[rowIndex]))
   result <- rbind(result, newRow)
  }
  colnames (result) <- c("timestamp", "outlier_value", "description")
  return (result)
 }
}

This function is just an elementary example. It is easy to add to each outlier more detailed information you can compile from the full data frame.

Time series with outliers at green markers

outliers with descriptions
  timestamp outlier_value description
1
2017-01-17 06:53:00
209
gear display flashing
2
2017-09-19 09:10:00
206
gear shift failure
3
2017-11-17 07:26:00
211
check engine lamp on

Dates are a sore point of analytics: they alway get you. When no time zone is specified, i.e., tz = "", R assumes the local time zone. In the data frame returned by Twitter's AnomalyDetectionTs functions, the time column has UTC as the time zone. Therefore, the following statement is useful after the call to AnomalyDetectionTs:

anomalies$anoms$timestamp <- as.POSIXct(anomalies$anoms$timestamp, tz = "")

Monday, April 24, 2017

Juggling Tools

Discussions about imaging invariably mention imaging pipelines. A simple pipeline to transform the image data to a different color space may have three stages: a lookup table to linearize the signal, a linear approximation to the second color space, and a lookup table to model the non-linearity of the target space. As an imaging product evolves, engineers add more pipeline stages: tone correction, gamut mapping, anti-aliasing, de-noising, sharpening, blurring, etc.

In the early days of digital image processing, researchers quickly realized that imaging pipelines should be considered harmful because, due to discretization, at each stage, the resulting image space became increasingly sparse. However, in the early 1990s, with the early digital cameras and consumer color printers, imaging pipelines came back. After some 25 years of experience, engineers have become more careful with the pipelines, but they are still a trap.

In data analytics, people often make a similar mistake. There are also three basic steps, namely data wrangling, statistical analysis, and presentation of the result. As development progresses, the analysis becomes richer; when the data is a signal, it is filtered in various ways to create different views, statistical analyses are applied, the data is modeled, classifiers are deployed, estimates and inferences are computed, etc. Each step is often considered as a separate task, encapsulated in a script that parses in a comma separated values (CSV) data file, calls one or more functions, and the writes out a new CSV file for the next stage.

The pipeline is not a good model to use when architecting a complex data processing endeavor.

I cannot remember if it was 1976 or 1978 when at PARC the design of the Dorado was finished and Chuck Thacker hand-wrote the first formal note on the next workstation: the Dragon. While the Dorado had a bit-sliced processor in ECL technology, the Dragon was designed as a multi-processor full-custom VLSI system in nMOS technology.

The design was much more complex than any chip design that had been previously attempted, especially after the underlying technology was switched from nMOS to CMOS. It became immediately evident that it was necessary to design new design automation (DA) tools that could handle such big VLSI chips.

A system based on full-custom VLSI design was a sequence of iterations of the following steps: design a circuit as a schematic, lay out the symbolic circuit geometry, check the design rules, perform logic and timing analysis, create a MOSIS tape, debug the chip. Using stepwise refinement, the process was repeated at the cadence of MOSIS runs. In reality, the process was very messy, because, at the same time, the physicists were working on the CMOS fab, the designers were creating the layout, the DA people were writing the tools, and the system people were porting the Cedar operating system. Just in the Computer Science Laboratory alone, about 50 scientists were working on the Dragon project.

The design rule checker Spinifex played a somewhat critical role, because it parsed the layout created with ChipNDale, analyzed the geometry, flagged the design rule errors, and generated the various input files for the logic simulator Rosemary and the timing simulator Thyme. Originally, Spinifex was an elegant hierarchical design rule checker, which allowed to verify all the geometry for a layout in memory. However, with the transition from nMOS to CMOS, the designers transitioned more and more to a partially flat design, which broke Spinifex. The situation was exacerbated by the endless negotiations between designers and physicists to allow for exceptions to the rules, leading to a number of complementary specialized design rule checkers.

With 50 scientists on the project, ChipNDale, Rosemary, and Thyme were also rapidly evolving. With the time pressure of the tape-outs, there were often inconsistencies in the various parsers. As the whipping boy in the middle of all this, one morning, while showering, I had an idea. The concept of a pipeline was contra naturam compared to the work process. The Smalltalk researchers on the other end of the building had an implementation process where a tree structure described some gestalt and methods would be written that decorate this representation of the gestalt.

In the following meeting, I proposed to define a data structure representing a chip. Tools like the circuit designer, the layout design tool, and the routers would add to the structure while tools like the design rule checkers and simulators would analyze the structure, with their output being further decorations added to the data structure. Even the documentation tools could be integrated. I did not expect this to have any consequence, but there were some very smart researchers in the room. Bertrand Serlet and Rick Barth implemented this paradigm and project representation and called it Core.

The power was immediately manifest. Everybody chipped in: Christian Jacobi, Christian Le Cocq, Pradeep Sindhu, Louis Monier, Mike Spreitzer and others joined Bertrand and Rick in rewriting the entire tool set around Core. Bob Hagman wrote the Summoner, which summoned all Dorados at PARC and dispatched parallel builds.

Core became an incredible game changer. While before there was never an entirely consistent system, now we could do nightly builds of the tools and the chips. Besides, the tools were no longer broken at the interfaces all the time.

The lubricant of the Silicon Valley are the brains wandering from one company to the other. When one brain wandered to the other side of the Coyote Hill, the core concept gradually became an important architectural paradigm that is on the basis of some modern operating systems.

If you are a data scientist, do not think in terms of scripts for pipelines connected by CSV files. Think of a core structure representing your data and the problem you are trying to solve. Think about literate programs that decorate your core structure. When you make the core structure persistent, think rich metadata and databases, not files with plain tables. Last but not least, also your report should be generated automatically by the system.

data + structure = knowledge

Thursday, April 13, 2017

Computational Imaging for Robust Sensing and Vision

In the early days of digital imaging, we were excited about having the images in numerical form and not being bound by the laws of physics. We had big ideas and quickly ran for their realization. However, we immediately reached the boundaries of the digital world: the computers of the day were too slow to process images, did not have enough memory, and the I/O was inadequate (from limited sensors to non-existing color printers).

Now has finally come the time when these dreams can be realized and computational color imaging has become possible, thanks to good sensors and displays, and racks full of general purpose graphical processing units (GPGPUs) with hundred of gigabytes of primary memory and petabytes of secondary storage. All this, at an affordable price.

Wednesday, 12 April 2017, Felix Heide gave a talk at The Stanford Center for Image Systems Engineering (SCIEN) with the title Capturing the “Invisible”: Computational Imaging for Robust Sensing and Vision. He presented three implementations.

One application is image classification. In the last couple of years we have seen what is possible with deep learning when you have a big Hadoop server farm and millions of users who provide large data sets they carefully label, creating gigantic training sets for machine learning. Felix Heide uses Bayesian inference to implement a much better system that is robust and fast. It better leverages the available ground-truth and uses proximal optimization to reduce the computational cost.

To facilitate the development of new algorithms, Felix Heide has created the ProxImaL Python-embedded modeling language for image optimization problems, available from www.proximal-lang.org.

computational imaging

Thursday, February 9, 2017

mirror mirror on the wall

Last November, I mentioned an app that makes you look like you are wearing a makeup when you do a teleconference. Now Panasonic lets you take it a step further. A new mirror analyzes the skin on your face and prints out a makeup that you can apply directly to your face.

The aim of the Snow Beauty Mirror is “to let people become what they want to be,” said Panasonic’s Sachiko Kawaguchi, who is in charge of the product’s development. “Since 2012 or 2013, many female high school students have taken advantage of blogs and other platforms to spread their own messages,” Kawaguchi said. “Now the trend is that, in this digital era, they change their faces (on a photo) as they like to make them appear as they want to be.”

When one sits in front of the computerized mirror, a camera and sensors start scanning the face to check the skin. It then shines a light to analyze reflection and absorption rates, find flaws like dark spots, wrinkles, and large pores, and offer tips on how to improve appearances.

But this is when the real “magic” begins. Tap print on the results screen and a special printer for the mirror churns out an ultrathin, 100-nanometer makeup-coated patch that is tailor-made for the person examined. The patch is made of a safe material often used for surgery so it can be directly applied to the face. Once the patch settles, it is barely noticeable and resists falling off unless sprayed with water.

The technologies behind the patch involve Panasonic’s know-how in organic light-emitting diodes (OLED), Kawaguchi said. By using the company’s technology to spray OLED material precisely onto display substrates, the printer connected to the computerized mirror prints a makeup ink that is made of material similar to that used in foundation, she added.

Read the full article by Shusuke Murai in the Japan Times News.

Panasonic Corp. engineer Masayo Fuchigami displays an ultrathin makeup patch during a demonstration of the Snow Beauty Mirror

Panasonic Corp. engineer Masayo Fuchigami displays an ultrathin makeup patch during a demonstration of the Snow Beauty Mirror on Dec. 1 in Tokyo. | Shusuke Murai

Thursday, January 19, 2017

Unable to complete backup. An error occurred while creating the backup folder

For the past four years, I have been backing up my laptop on a G-Technology Firewire disk connected to the hub in my display. So far it worked without a hitch, but a few days ago I started to get the error message

Time Machine couldn’t complete the backup to “hikae”.
Unable to complete backup. An error occurred while creating the backup folder.

The message appeared without a time pattern, so it was not clear what it could be. The drive could not be unmounted and had to be force-ejected and power-cycled and then worked again until the next irregular event, maybe one backup out of ten.

When I ran Disk Utility to see if something was wrong with the drive, it told me the boot block was corrupted. After fixing it, the Time Machine problem did not go away, so I must have corrupted the boot block with the force-eject. Time to find out what is going on.

The next time it happened, I tried to eject the drive from Disk Utility, which gave me the message

Disk cannot be unmounted because it is in use.

Who on Earth would be using it? Did Time Machine hang? Unix to the rescue, let us get the list of open files

sudo lsof /Volumes/hikae

The user is root and the commands are mds and mds_store on index files. They are indexing the drive for Spotlight. Why on Earth would an operating system index a backup drive by default? Let us get rid of that.

sudo mdutil -i off /Volumes/hikae

However, in this state, the command returns "Error: unable to perform operation. (-400) Error: unknown indexing state." This might mean Spotlight has crashed or is otherwise hanging.

Force Eject and power cycle the drive. This time mdutil works:

/Volumes/hikae:
2017-01-18 17:10:00.657 mdutil[25737:7707511] mdutil disabling Spotlight: /Volumes/hikae -> kMDConfigSearchLevelFSSearchOnly\\Indexing and searching disabled.

For the past two days, I have no longer experienced the problem.

If you are the product manager, why is Spotlight indexing backup drives by default?

If you prefer using a GUI, drag and drop your backup drive icon into the privacy pane of the Spotlight preference window (I did not try this):

Tell Spotlight not to index your backup drive

Friday, November 11, 2016

App adds makeup to faces on video conferences

In a potential boost for the government’s drive to get more people telecommuting, cosmetics company Shiseido Co. has developed an app that makes users look as if they are wearing makeup. It amounts to an instant makeover for the unfortunate worker called to appear on screen from home at an awkward hour.

Read the article in the Japan Times.

Yoko's lips

Thursday, October 13, 2016

Dataset metadata for search engine optimization

Last week I wrote a post on metadata. Google is experimenting with a new metadata schema it calls Science Datasets that will allow it to better make public datasets discoverable.

The mechanism is under development and they are currently soliciting interested parties with the following kinds of public data:

  • A table or a CSV file with some data
  • A file in a proprietary format that contains data
  • A collection of files that together constitute some meaningful dataset
  • A structured object with data in some other format that you might want to load into a special tool for processing
  • Images capturing the data
  • Anything that looks like a dataset to you

In your metadata schema you can use any of the schema.org dataset properties, but it should contain at least the following basic properties: name, description, url, sameAs, version, keywords, variableMeasured, and creator.name. If your dataset is part of a corpus, you can reference it in the includedInDataCatalog property.

There are also properties for download information, temporal coverage, spatial coverage, citations and publications, and provenance and license information.

This is a worthwhile effort to make your research and public datasets more useful to the community.

Creative Commons LicenseGoogle

Tuesday, October 4, 2016

Metadata

As Carlsson notes, big data in not about "big" but about complexity in format and structure. We can approach the format complexity through metadata, which allows us to navigate through the data sets and to determine what they are about.

Two important requirements on experiments are replicability and reproducibility. Replicability refers to the ability to rerun the exact data experiment to produce exactly the same result; it is an aspect of governance and it is good practice to always have somebody else to check the data and its analysis before it is published. Reproducibility refers to the ability to use different data, techniques, and equipment to confirm the same result as previously obtained. We can be confident in a result only after it has been reproduced independently. These two requirements guide us to what kind of metadata we need.

There are three classes of metadata: context, syntax, and semantic.

Context of data refers to how, when, and where it was collected. The context is usually written in a lab book. If we need to replicate an analysis at a later time, the lab book might be unretrievable, therefore the context of data has to be stored with the data. This can also be a big money and time saver because some ancillary data we need for an analysis might already be available from a previous experiment; we need to be able to find it.

The syntax of data refers to the format. Analysts spend a large amount of their time wrangling data. When the format of each time series is clearly described, this tedious work can be greatly simplified. During replication and reproduction, it can also help diagnose such frequent errors like the confusion between metric and imperial units of measure. Ideally, with the data, we should also store the APIs to the data because they are part of the syntax of data.

The semantic of data refers to its meaning and is the most difficult metadata to produce. We require a unified framework that researchers in all scientific disciplines can use to create consistent, easily searchable metadata. Ease of use is paramount. Because the ability to share data is so important, we want the process of metadata creation to be as painless as possible. This means that we must start by creating an ontology for each domain in which we create data.

Ontologies evolve with time. a big challenge is to track this evolution with the metadata. For example, if we called a technique "machine learning" but then realize the term is too generic and we should call it "cluster analysis" because this is what we were doing anyway, we have to update also the old metadata. Data curation applies also to the metadata.

the evolution of terms

Some metadata can be computed from the data itself, for example, the descriptive statistics. At NASA, the automatic extraction of metadata from data content is called data archeology.

Friday, September 30, 2016

Navigating instead of searching

I believe it was on a day at the end of March or beginning of April 1996, when out of the blue I received the assignment to write a report on why the world wide web would be important for my employer at the time. I was given only two weeks to write a blurb and a slide deck. I had not thought about this matter and it was a struggle to write something meaningful in such a short time, without any time to read up. It ended up as an exercise in automatic writing: just write ahead and never look back and revise.

At the end, I delivered the blurb, but management decided the web was just a short-lived fad like citizen band (CB) radio that would go away shortly and put the blurb in the technical report queue for external publication. I did not think much about it because such is life in a research lab. However, with all the hype of the time on the "Internet Tsunami" by Bill Gates, the "Dot in the Dot Com" by Scott McNealy, and all the others—while my employer remained silent and kept everybody curious—the report W3 + Structure = Knowledge was requested hundreds of times (at that time a tech report was typically requested much less than ten times). Subsequently, I received quite a few requests to present the companion slide deck.

In my struggle to write something in two weeks, I typed about the need to structure information on the world wide web so it can be easily navigated, instead of searching for information.

little dancer leaping over the world wide web

It appears today we are at such a disruptive juncture again. This time, it is not about websites: it is about data (some prefer to call it big data, but size does not really matter). Solid state drives are now inexpensive, can hold 60 TB of data in each 3.5" drive, and have access times similar to RAM. In addition, we have all the shared data in the various clouds.

Today, we are not interested in finding data, we are maybe wiling to navigate to data, but preferably we would like the data to anticipate our need for it and come to us in digested, actionable form. Actually, we are not interested in the data, we are interested in data in a context: knowledge. We want the data to come to us and ask us if it is OK to take an anticipated action based on a compiled body of knowledge: wisdom.

This is an emergent property because all the pieces have fallen together. Our mobile devices and the internet of things constantly gather data. There are open source implementations of deep learning algorithms. CUDA 8 lets us run them on inexpensive Pascal GPGPUs like the Tesla P100. Algebraic topology analytics lets us build networks that compile knowledge about the data. Digital assistant technology brings this knowledge at our service.

Another key ingredient is the skilled workforce. Mostly Google, but also Facebook, Apple, Amazon, etc. have been aggressively educating their workforce in advanced analytics, and as brains move from company to company in the Silicon Valley, these skills are diffused in the industry.

In a recent New York Times article, G.E.'s Jeffrey R. Immelt explained how he is taking advantage of this talent pool in a new Silicon Valley R&D facility employing 1,400 people. Microsoft is creating a new group, the AI and Research Group, by combining the existing Microsoft Research group with the Bing and Cortana product groups, along with the teams working on ambient computing. Together, the new AI and Research Group will have some 5,000 engineers and computer scientists.

This is the end of search engines. This is the end of metadata: we want wisdom based on all the data.

Here is a revised version of the post I wrote in July.

The asumption is that to be useful, technology has to enable society to become more efficient so life quality increases. The increase has to be at least one order of magnitude.

Structured data

In the context of big data, we read a lot about structured versus unstructured data. So far, so good. Things get a little murky and confusing when advanced analytics—which refers to analytics for big data—joins the conversation. The confusion comes from the subtle difference between "structured data" and "structure of data," which contain almost the same words (their edit distance is 3). Both concepts are key to advanced analytics, so they often come up together. In this post, we will try to shed some light on this murkiness to clarify it.

The categorization in structured, semi-structured, and unstructured data comes from the storage industry. Computers are good at chewing on large amounts of data of the same kind, like for example the readings from a meter or sensor, or the transactions on cash registers. The data is structured in the sense that each record has the same fields at the same locations, for example on an 80 or 96 column punched card, if you want a visual image. This structure is described in a schema.

Databases are optimized for storing structured data. Since each record has the same structure, the location of the i-th record on the disk is i times the record length. Therefore, it is not necessary to have a file system: a simple block storage system is all that is needed. When instead of the i-th record we need the record containing a given value in a given field, we have to scan the entire database. When this is a frequent operation in a batch step, we can accelerate it by first sorting the records by the values in this field, which allows us to use binary search, which is logarithmic instead of linear.

Because an important performance metric is the number of transactions per second, database management systems use auxiliary structures like index files and optimizing query systems like SQL. In a server-based system, when we have a query, we do not want to transfer the database record by record to the client: this leads to server-based queries. When there are many clients, often the same query is issued from various clients, therefore, caching is an important mechanism to optimize the number of transactions per second.

Database management systems are very good at dealing with transactions on structured data. There are many optimization points that allow for huge performance gains, but it is a difficult art requiring highly specialized analysts.

Semi-structured data

With cloud computing, it has become very easy to quickly deploy a consumer application. The differentiation is no longer by the optimization of the database, but in being able to collect and aggregate user data so it can be sold. This process is known as monetization and an example is click-streams. The data is to a large extent in the form of logs, but their structure is often unknown. One reason is that the schemata often change without a notification because the monetizers infer them by reverse engineering. Since the data is structured with an unknown schema, it is called semi-structured. With the Internet of Things (IoT), also known as Web of Things, Industrial Internet, etc., a massive source of semi-structured data is coming towards us.

This semi-structured data is high-volume and high-velocity. This breaks traditional relational databases because data parsing and schema inference become a performance bottleneck. Also, the indexing facilities may not be able to cope with the data volume. Finally, the traditional database vendor's pricing models do not work for high volumes of less costly data. The paradigms for semi-structured data are column based storage and NoSQL (not only SQL).

In big data scenarios, structured data can have high-volume and high-velocity. Although it may be fully structured, e.g., rows of double precision floating point values from a set of sensors (a time series), a commercial database system might lose data when reindexing. Even a NoSQL database might be too slow. In this case, this structured data is treated as unstructured and each column is stored in a separate file for concurrent writes. Typically, the content of such a file is a time series.

Unstructured data

The ubiquity of smartphones with their photo and video capabilities and connectedness to the cloud has brought a flood of large data files. For example, when the consumer insurance industry thought it can streamline its operations by having insured customers upload images of damages instead of keeping a large number of claim adjusters in the field, they got flooded with images. While an adjuster knows how to document a damage with a few photographs, consumers take dozens of images because they do not know what is essential.

Photographs and videos have a variety of image dimensions, resolutions, compression factors, and duration. The file sizes vary from a few dozen kilobytes to gigabytes. They cannot be stored in a database other than as a blob, for binary large object: the multimedia item is stored as a file or an object and the database just contains a file pathname or the address of an object. In general, not just images and video are stored in blobs, therefore we use the more generic term of digital items. Examples of digital items in engineering applications are drawings, simulations, and documentation. In their 2006 paper, Jim Gray et al. found that databases can store efficiently digital items of up to 256 KB [1].

For digital items, in juxtaposition to conventional structured data, the storage industry talks about unstructured data.

Unstructured data can be stored and retrieved, but there is nothing else that can be done with it when we just look at it as a blob. When we look at analytics jobs, we see that analysts spend most of their time munging and wrangling data. This task is nothing else than structuring data because analytics is applied to structured data.

In the case of time series data, the wrangling is easy, as long as the columns have the same length. If not, a timestamp is needed to align the time series elements and introduce NA values where data is missing. An example of misaligned data is when data from various sources is blended.

In the case of semi-structured data, wrangling entails reverse engineering the schema, convert dates between formats, distinguish numbers and strings from factors, and dealing correctly with missing data. In the case of unstructured data, it is about extracting the metadata by parsing the file. This can be a number of tags like the color space, or it can be a more complex data structure like the EXIF, IPTC, or XMP metadata.

Structure in time series

A time series is just a bunch of data points, so at first one might think there is no structure. In a way, statistics, with its aim to summarize data, can describe the structure in raw data. It can infer its distribution and its parameters, model it through regression, etc. These summary statistics are the emergent metadata of the time series.

Structure in images

A pictorial image is usually compressed with the JPEG method and stored in a JFIF file. The metadata in a JPEG image consists of segments beginning with a marker, the kind of the marker, and if there is a payload, the length and the payload itself. An example of a marker kind is the type (baseline or progressive) followed by width, height, number of components, and their subsampling. Other markers are the Huffman table (HT), the quantization tables (DQT), a comment, and application-specific markers like the color space, color gamut, etc. This illustrates that unstructured data contains a lot of structure. Once the data wrangler has extracted and munged this data, it is usually stored in R frames, or in a dedicated HIVE or MySQL database. These allow processing with analytics software.

Deeper structure in images

Analytics is about finding even deeper structure in the data. For example, a JPEG image is first partitioned in 8×8 pixel blocks, which are each subjected to a DCT. Pictorially, the cosine basis (the kernels) looks like in this figure:

the kernels of the discrete cosinus transform (DCT)

The DCT transforms the data into the frequency domain, similar to the discrete Fourier transform, but in the real domain. We do this to decorrelate the data. In each of the 64 dimensions, we determine the number of bits necessary to express the values without perceptual loss, i.e., in dependence of the MTF of the combined HVS and capture & viewing devices. These numbers of bits are what is stored in the discrete quantization table DQT, and we zero out the lower order bits, i.e., we quantize the values. At this point, we have not reduced the storage size of the image, but we have introduced many zeros. Now we can analyze statistically the bit patterns in the image representation and determine the optimal Huffman table, which is stored with the HT marker, and we compress the bits, reducing the storage size of the image through entropy coding.

Like we determine the optimal HT, we can also study the variation in the DCT-transformed image and optimize the DQT. Once we have implemented this code, we can use it for analytics. We can compute the energy in each of the 64 dimensions of the transformed image. As a proxy for energy, we can compute the variance and obtain a histogram with 64 abscissa points. The shape of ordinates gives us an indication of the content of the image. For example, the histogram will tell us, if an image is more likely scanned text or a landscape.

We have built a rudimentary classifier, which gives us a more detailed structural view of the image.

Let us recapitulate: we transform the image to a space that gives a representation with a better decorrelation (like transforming from RGB to CIELAB, then from Euclidean space to cosine space). Then we perform a quantization of the values and study the histogram of the energy in the 64 dimensions. We start with a number of known images and obtain a number of histogram shapes: this is the training phase. Then we can take a new image and estimate its class by looking at its DCT histogram: we have built a rudimentary classifier.

An intuition for deep learning

We have used the DCT. To generalize, we can build a pipeline of different transformations followed by quantizations. In the training phase, we determine how well the final histogram classifies the input and propagate back the result by adjusting the quantization table, i.e., to fine-tune the weights making up the table elements. In essence, this is an intuition for what happens in deep learning, or more formally, in CNN.

In the case of image classification and object recognition, the CNN equivalent of a JPEG block is called receptive field and the generalization of the DCT DQT is called CNN filter or CNN kernel. The CNN kernel elements are the weights, just as they are the coordinates in cosine space in the case of JPEG (this is where the analogy ends: while in JPEG the kernels are the basis elements of the cosine space, in CNN the kernels are the weights and the basis in unknown). The operation of applying each filter in parallel to each receptive field is a convolution. The result of the convolution is called the feature map. While in JPEG the kernels are the discrete cosine basis elements, in CNN the filters can be arbitrary feature identifiers. While in JPEG the kernels are applied to each block, in CNN they “slide” by a number of pixels called the stride (in JPEG this would be the same as the block size, in CNN it is always smaller than the receptive field).

The feature map is the input for a new filter set and the process is iterated. However, to obtain a faster convergence, after the convolution it is necessary to introduce a nonlinearity, in a so-called ReLU layer. Further, the feature map is down-sampled, using what in CNN is called a pooling layer: this helps to avoid to overfit the training set. At the end, we have the fully connected layer, where for each class being considered, there is the probability for the initial image to belong to that class:

CC Aphex34 (Wikimedia Commons)

The filters are determined by training. We start with a relatively large set of images for whom we have determined manually the probabilities to belong to the various classes: this is the ground truth and the process is called groundtruthing. The filters can be seeded with random patterns (in reality we use the filters from a similar classification problem: transfer learning or pre-training, see the figure below) and we apply the CNN algorithm to the ground truth images. At the end, we look at the computed connected layer and compare it with the ground truth obtaining the loss function. Now we propagate the errors back to reduce the weights in the filters that caused the largest losses. The process of forward pass, loss function back-propagation, and weight update is called an epoch. The training phase is applied over a large ground truth over many epochs.

Evolution of pre-trained conv-1 filters with time; after 5k, 15k, and 305k iterations, from [2]:

Evolution of pre-trained conv-1 filters with time; after 5k, 15k, and 305k iterations, from Agrawal et al.

CNN and all the following methods are only possible because of the recent progress in GPGPUs.

A disadvantage of machine learning is that the training phase takes a long time and if we change the kind of input we have to retrain the algorithm. For some applications, you can use your customers as free workers, like for OCR you can use the training set as captcha, which your customers will classify for free. For scientific and engineering applications, you typically do not have the required millions of free workers. This is the motivation for unsupervised machine learning.

Assimilation and accommodation

So far, we took a random walk from unstructured data to multimedia files, JPEG compression, a DCT-inspired classifier, and deep learning. We saw that the crux of supervised machine learning is the training.

There are two reasons for needing classifiers. We can design more precise algorithms if we can specialize them for a certain data class. For humans, the reason is that our immediate memory can hold only 7 ± 2 chunks of information. This means that we aim to break down information into categories each holding 7 ± 2 chunks. There is no way humans can interpret the graphical representation of graphs with billions of nodes.

As already Immanuel Kant noted, categories are not natural or genetic entities, they are purely the product of acquired knowledge. One of the functions of the school system is to create a common cultural background, so people learn to categorize according to similar rules and understand each other's classifications. For example, in the biology class, we learn to organize botany according to the 1735 Systema Naturæ compiled by Carl Linnæus.

As we know from Jean Piaget's epistemological studies with children, there is assimilation when a child responds to a new event in a way that is consistent with an existing classification schema. There is accommodation when a child either modifies an existing schema or forms an entirely new schema to deal with a new object or event. Piaget conceived intellectual development as an upward expanding spiral in which children must constantly reconstruct the ideas formed at earlier levels with new, higher order concepts acquired at the next level.

The data scientist's social role is to further expand this spiral. Concomitantly, data scientists have to be well aware of the cultural dependencies of acquired knowledge.

For data, this means that we want to cluster it (recoding by categorization). Further, we want to connect the clusters in a graph so we can understand its structure (finding patterns). At first, clustering looks easy: we take the training set and do a Delaunay triangulation, the dual graph of the Voronoi diagram. After building the graph with the training set, for a new data point, we just look in which triangle it falls and know its category. Color scientists are familiar with Delaunay triangulations because they are used for device modeling by table lookup. Engineers use them to build meshes for finite element methods.

The problem is that the data is statistical. There is no clear-cut triangulation and points from one category can lie in a nearby category with a certain probability. Roughly, we build clusters by taking neighborhoods around the points and the intersect them to form the clusters. The crux is to know what radius to pick for the neighborhoods because the outcome will be very different.

Algebraic topology analytics

This is where the relatively new field of algebraic topology analytics comes into play [4]. It has only been about 15 years that topology has started looking at point clouds. Topology, an idea of the Swiss mathematician Leonhard Euler, studies the properties of shape independent of coordinate systems, dependent only on a metric. The topological properties are deformation invariant (a donut is topologically equivalent to a mug). Finally, topology constructs compressed representations of shape.

The interesting element of shape in point clouds are the k-th Betti numbers βk, the number of k-dimensional "holes" in a simplicial complex. For example, informally β0 is the number of connected components, β1 the number of roundish holes, and β2 the number of cavities.

Algebraic topology analytics relieves the data scientist from having to guess the correct radius of the point neighborhoods by considering all radii and retaining only those that change the topology. If you want to visualize this idea, you can think of a dendrogram. You start with all the points and represent them as leaves; as the radii increase, you walk up the hierarchy in the dendrogram.

This solves the issue of having to guess a good radius to form the clusters, but you still have the crux of having to find the most suitable distance metric for your data set. This framework is not a dumb black-box: you still need the skills and experience of a data scientist.

The dendrogram is not sufficiently powerful to describe the shape of point clouds. The better tool is the set of k-dimensional persistence barcodes that show the Betti numbers in function of the neighborhood radii for building the simplicial complexes. Here is an example from page 347 in Carlsson's article [4]:

(a) Zero-dimensional, (b) one-dimensional, and (c) two-dimensional persistence barcodes

Motifs

With large data sets, when we have a graph, we do not necessarily have something we can look at because there is too much information. Often we have small patterns or motifs and we want to study how a higher order graph is captured by a motif [5]. This is also a clustering framework.

For example, we can look at the Stanford web graph at some time in 2002 when there were 281,903 nodes (pages) and 2,312,497 edges (links).

Clusters in the Stanford web graph

We want to find the core group of nodes with many incoming links and the tied together periphery groups that are tied together and also up-link to the core.

A motif that works well for social network kind of data is that of three interlinked nodes. Here are the motifs with three nodes and three edges:

Motifs for social networks

In motif M7 we marked the top node in red to match the figure of the Stanford web.

Conceptually, given a higher order graph and a motif Mi, the framework searches for a cluster of nodes S with two goals:

  1. the nodes in S should participate in many instances of Mi
  2. the set S should avoid cutting instances of Mi, which occurs when only a subset of the nodes from a motif are in the set S

The mathematical basis for this framework are motif adjacency matrices and the motif Laplacian. With these tools, a conductance metric in spectral graph theory can be defined, which is minimized to find S. The paper [5] in the references below contains several worked through examples for those who want to understand the framework.

References

  1. R. Sears, C. van Ingen, and J. Gray. To BLOB or not to BLOB: Large object storage in a database or a filesystem? Technical Report MSR-TR-2006-45, Microsoft Research, One Microsoft Way, Redmond, WA 98052, April 2006.
  2. P. Agrawal, R. Girshick, and J. Malik. Analyzing the Performance of Multilayer Neural Networks for Object Recognition, pages 329–344. Springer International Publishing, Cham, September 2014.
  3. G. A. Miller. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2):81–97, 1956
  4. G. Carlsson. Topological pattern recognition for point cloud data. Acta Numerica, 23:289–368, 5 2014
  5. A. R. Benson, D. F. Gleich, and J. Leskovec. Higher-order organization of complex networks. Science, 353(6295):163–166, 2016

Thursday, August 18, 2016

Typesetting Sweave documents with bibliographies

When you just have to make a quick plot, you can script R on the console, but when you do an experiment and want to be able to reproduce it, scripts are no good. I architect a solution and then use an IDE like RStudio to write a program. Once I have the wrangling and helper functions working the way I need, I copy them from their R files into a Sweave template and continue from there.

Today I got stuck because I was using a bibliography and RStudio could not typeset it.

It turns out, that RStudio just appears to run the LaTeX typesetting command, not the pdflatexmk script, so Biber is not called. When you want to work fast and just point and click instead of typing on the command line, you double-click on the tex file and typeset it in TeXShop.

Unfortunately, that does not work out-of-the-box because the Sweave package is part of the R distribution but not the TeXLive distribution. The solution is to make a copy of the two Sweave style files in your local TeXLive texmf directory. Here are the four steps for the current MacOS versions:

  1. go to /Library/Frameworks/R.framework/Resources/share/texmf/tex/latex
  2. copy Rd.sty and Sweave.sty
  3. go to ~/Library/texmf/tex/latex and paste the two files
  4. do a sudo texhash

Now you can typeset your Sweave documents both in Studio and TeXShop. The latter is handy only when you need to redo the bibliography, glossary, or index: you can keep working just in RStudio.

The very first line in your LaTeX file, before the class declaration, should be

% !TEX TS-program = pdflatexmk

Quantile-Quantile Plot

Tuesday, August 9, 2016

A new open forum for scientists working on color

The American Association for the Advancement of Science (AAAS) is setting up a new platform for scientific collaboration called Trellis. A key feature is that you can upload papers you want to discuss and anybody in the group can read the paper online: the AAAS takes care of all the copyright issues with the journal publisher. Another feature is that you can control how much noise you get from Trellis. Messages can have up to 2500 characters, but if you have longer text you can create a PDF and upload it like a paper. The same holds for images and videos.

Many AAAS members are teachers: if you do crowd-sourced experiments, you can easily find subjects. The AAAS is also involved in policy making in Washington, in case you need help with that. You can also announce conferences and other scientific events, and use the shared calendar.

Trellis is currently in pilot phase for educators, policy makers, and Section T of AAAS (Information, Computing, and Communication). There are still a few rough edges that are being ironed out. That might be why you have not yet heard of Trellis.

I am setting up a group called Computational Color Science in Section T. However, we are trying out something new. We are making it a completely open group, i.e., anybody with the URL can sign up and participate, without having to be an AAAS member. To join, go to http://www.trelliscience.com/color/. You can invite anybody else you want by giving them the URL and they can sign up.

Tuesday, July 19, 2016

Structure in Unstructured Data, Part 2

Follow this link for an updated post.

In the first part, we took a random walk from unstructured data to multimedia files, JPEG compression, a DCT-inspired classifier, and deep learning. We saw that the crux of supervised machine learning is the training.

There are two reasons for needing classifiers. We can design more precise algorithms if we can specialize them for a certain data class. For humans, the reason is that our immediate memory can hold only 7±2 chunks of information. This means that we aim to break down information into categories each holding 7±2 chunks. There is no way humans can interpret the graphical representation of graphs with billions of nodes.

As already Immanuel Kant noted, categories are not natural or genetic entities, they are purely the product of acquired knowledge. One of the functions of the school system is to create a common cultural background, so people learn to categorize according to similar rules and understand each other's classifications. For example, in the biology class, we learn to organize botany according to the 1735 Systema Naturæ compiled by Carl Linnæus.

As we know from Jean Piaget's epistemological studies with children, there is assimilation when a child responds to a new event in a way that is consistent with an existing classification schema. There is accommodation when a child either modifies an existing schema or forms an entirely new schema to deal with a new object or event. Piaget conceived intellectual development as an upward expanding spiral in which children must constantly reconstruct the ideas formed at earlier levels with new, higher order concepts acquired at the next level.

The data scientist's social role is to further expand this spiral.

For data, this means that we want to cluster it (recoding by categorization). Further, we want to connect the clusters in a graph so we can understand its structure (finding patterns). At first, clustering looks easy: we take the training set and do a Delaunay triangulation, the dual graph of the Voronoi diagram. After building the graph with the training set, for a new data point, we just look in which triangle it falls and know its category. Color scientists are familiar with Delaunay triangulations because they are used for device modeling by table lookup. Engineers use them to build meshes for finite element methods.

The problem is that the data is statistical. There is no clear-cut triangulation and points from one category can lie in a nearby category with a certain probability. Roughly, we build clusters by taking neighborhoods around the points and the intersect them to form the clusters. The crux is to know what radius to pick for the neighborhoods because the result will be very different.

This is where the relatively new field of algebraic topology analytics comes into play. It has only been about 15 years that topology has started looking at point clouds. Topology, an idea of the Swiss mathematician Leonhard Euler, studies the properties of shape independent of coordinate systems, dependent only on a metric. The topological properties are deformation invariant (a donut is topologically equivalent to a mug). Finally, topology constructs compressed representations of shape.

The interesting element of shape in point clouds are the k-th Betti numbers βk, the number of k-dimensional "holes" in a simplicial complex. For example, informally β0 is the number of connected components, β1 the number of roundish holes, and β2 the number of cavities.

Algebraic topology analytics relieves the data scientist from having to guess the correct radius of the point neighborhoods by considering all radii and retaining only those that change the topology. If you want to visualize this idea, you can think of a dendrogram. You start with all the points and represent them as leaves; as the radii increase, you walk up the hierarchy in the dendrogram.

This solves the issue of having to guess a good radius to form the clusters, but you still have the crux of having to find the most suitable distance metric for your data set. This framework is not a dumb black-box: you still need the skills and experience of a data scientist.

The dendrogram is not sufficiently powerful to describe the shape of point clouds. The better tool is the set of k-dimensional persistence barcodes that show the Betti numbers in function of the neighborhood radii for building the simplicial complexes. Here is an example from page 347 in Carlsson's article cited below:

(a) Zero-dimensional, (b) one-dimensional, and (c) two-dimensional persistence barcodes

With large data sets, when we have a graph, we do not necessarily have something we can look at because there is too much information. Often we have small patterns or motifs and we want to study how a higher order graph is captured by a motif. This is also a clustering framework.

For example, we can look at the Stanford web graph at some time in 2002 when there were 281,903 nodes (pages) and 2,312,497 edges (links).

Clusters in the Stanford web graph

We want to find the core group of nodes with many incoming links and the tied together periphery groups that are tied together and also up-link to the core.

A motif that works well for social network kind of data is that of three interlinked nodes. Here are the motifs with three nodes and three edges:

Motifs for social networks

In motif M7 we marked the top node in red to match the figure of the Stanford web.

Conceptually, given a higher order graph and a motif Mi, the framework searches for a cluster of nodes S with two goals:

  1. the nodes in S should participate in many instances of Mi
  2. the set S should avoid cutting instances of Mi, which occurs when only a subset of the nodes from a motif are in the set S

The mathematical basis for this framework are motif adjacency matrices and the motif Laplacian. With these tools, a conductance metric in spectral graph theory can be defined, which is minimized to find S. The third paper in the references below contains several worked through examples for those who want to understand the framework.

Further reading: