Follow the exploration on how to build a better taxonomy for business. I believe current taxonomies are currently lacking. Using a spatial concept of taxonomy, our team has built a taxonomy that is used to power research into Comparables for Mergers & Acquisitions in the evolving sector of Information, Software and Media, and I would like to share our thoughts on how to build a better taxonomy.
Showing posts with label Classification. Show all posts
Showing posts with label Classification. Show all posts
Thursday, January 5, 2012
Taxonomy Mapping Engine in Use
Our team is in Annual Trends Report Mode. Yesterday we published our first of seven Trends Reports tracking Mergers & Acquisitions in the Media and Software Sectors. This report is for the media space. One part of the report is to collect the deals that for this space. We use an auto-population algorithm that fills the lists many times a day. We then use an Industry Map and rules set that maps the categorized deals into a simple flat taxonomy just used for this report. We can then compare sub-segments of the Sector to see which segments are performing better or worse. Note we do this for 7 different reports each with its own Industry Map and rules set. We never have to categorize a deal more than one time. The mapping engine puts the deals into the appropriate bins for that report. Check it out here.
Wednesday, January 4, 2012
Flying a Kite - Starting a Taxonomy
For Christmas, I got this wonderful book about the Brooklyn Bridge. It is called "The Great Bridge" by David McCullough. One part, he wrote about the first bridge to cross the Niagara Gorge built by Charles Ellet. The way Ellet started the bridge was to offer five dollars the first American boy who could fly a kite over to the Canadian side of the gorge. The bridge span was 1,010 feet, and young Homer Walsh won the prize. Ellet took the kite string that spanned the gorge, and tied successively heavier cords and pulled them across the gorge until he had a heavy cable spanning the gorge and from that he built his bridge.
This story reminded me of our team's first efforts of building a business taxonomy. We started with a simple flat set of categories, and then added a second set of categories. After that we migrated to hierarchical trees, and then to banyans, and today we are ever adding features and complexity to our business taxonomy. But we could not have gotten to where we are today, unless we had first tried our first simple solution to span our own problem.
This story reminded me of our team's first efforts of building a business taxonomy. We started with a simple flat set of categories, and then added a second set of categories. After that we migrated to hierarchical trees, and then to banyans, and today we are ever adding features and complexity to our business taxonomy. But we could not have gotten to where we are today, unless we had first tried our first simple solution to span our own problem.
Wednesday, December 21, 2011
Challenges of Classifying a Business
One of the biggest challenges our team has building a taxonomy of businesses is the actual classifying of a business. The process at best is semi-automated. Our classifiers generally have to look at the website of a given business and try to determine what they do. The website, generally speaking, is not written to describe how a business operates, but rather to sell the business's products and services. Each website has its own content, and we used to concentrate on the "About Us" page where the business defines itself. Unfortunately, the "About Us" page is usually some sort of vague mission statement. We even looked at our own website, and found the wording written by our Marketing guy was so general that you would not know what we did! So our classifiers have learned to scan around many pages of a businesses site to learn what they do, how they do it and who they sell to. We have built some scrapers, but our results have been mediocre. We are now currently looking at ways to scrape company websites and intelligently gather info data to further automate the classification. The big problem we saw was that matching words on the website against words in our thesaurus gave us too many false positives. We are now looking at weighting the value of words in our thesaurus and how they matched in the past with verified classifications. Anyone explore these types of auto-classification.
Wednesday, December 14, 2011
Banyan provides an interesting view on our data.
The folks at Ongig.com, a new video enhanced job site, asked for some of our data to see who are the top 25 most efficient Tech companies based on profit per employee. The winner is SanDisk followed by Google and Apple.
http://ongig.com/blog/wall-street/profit-per-employee-tech#more-1399
We were able to easily provide this because our chief taxonomist had created a top level node in our industry tree called Information Technology which has child nodes for Computer Manufacturers, Software Companies, Electronic Information, and IT Consultants. All these nodes have other parents elsewhere in the tree, but the Banyan helps pull it all together.
I would also like to add a comment by Leo Meerman from LinkedIn. He mentioned that what I call a "Banyan Tree" is better known as a "Polyhierarchy".
http://ongig.com/blog/wall-street/profit-per-employee-tech#more-1399
We were able to easily provide this because our chief taxonomist had created a top level node in our industry tree called Information Technology which has child nodes for Computer Manufacturers, Software Companies, Electronic Information, and IT Consultants. All these nodes have other parents elsewhere in the tree, but the Banyan helps pull it all together.
I would also like to add a comment by Leo Meerman from LinkedIn. He mentioned that what I call a "Banyan Tree" is better known as a "Polyhierarchy".
Tuesday, December 13, 2011
The Banyan Tree - a new hierarchy
Most taxonomies are set up in a hierarchical tree format. Our team consists of four distinct trees for each facet of how a business operates. Developing the taxonomy and software over the last 8 years, we found at some point that the tree structure became too strict. There were certain categories that did not want to be under just one parent category. A prime example of this situation is video game companies. These companies make software for entertainment purposes. As this industry has matured, it has become closely link with the big entertainment companies, and they employ teams of artists, writers as well as programmers. Historically, these businesses should be software, but they are also so tied closely to entertainment companies it seems odd that when searching for entertainment companies that these would not turn up in our search results. One solution is to move video game studios to under the entertainment category, but then we lose the software aspect of the business. Our team's solution was to change the nature of our trees. We now allow nodes to have multiple parents, and create what I call the "Banyan" tree, which is a tree from India that has multiple trunks to the ground.
Looking at the Wikipedia article, we see that the banyan tree name comes from the Gujarti word for merchant, because merchant markets were often located under these great trees. It seems appropriate for a business taxonomy.
Wednesday, November 23, 2011
Why the North American Industry Classification System (NAICS) is no good?
When navigating a database of businesses, you need a taxonomy in order to find companies in an industry you are interested in. You would think that the NAICS would be ideal, however in practice none of the commercial databases use it. The reason is found on the US Census web site.
As stated on US Census web site, "The North American Industry Classification System (NAICS) is the standard used by Federal statistical agencies in classifying business establishments for the purpose of collecting, analyzing, and publishing statistical data related to the U.S. business economy." and "NAICS was developed under the auspices of the Office of Management and Budget (OMB), and adopted in 1997 to replace the Standard Industrial Classification (SIC) system. It was developed jointly by the U.S. Economic Classification Policy Committee (ECPC), Statistics Canada
, and Mexico's Instituto Nacional de Estadistica y Geografia
, to allow for a high level of comparability in business statistics among the North American countries."
The reason that it is not useful is that it is used to track broad trends. When you need to analyze business segments of our market you will see that a finer grained and richer taxonomy is needed. I have started this blog to explore this issue as my team and I continue to tackle the issues. Stay tuned.
As stated on US Census web site, "The North American Industry Classification System (NAICS) is the standard used by Federal statistical agencies in classifying business establishments for the purpose of collecting, analyzing, and publishing statistical data related to the U.S. business economy." and "NAICS was developed under the auspices of the Office of Management and Budget (OMB), and adopted in 1997 to replace the Standard Industrial Classification (SIC) system. It was developed jointly by the U.S. Economic Classification Policy Committee (ECPC), Statistics Canada
The reason that it is not useful is that it is used to track broad trends. When you need to analyze business segments of our market you will see that a finer grained and richer taxonomy is needed. I have started this blog to explore this issue as my team and I continue to tackle the issues. Stay tuned.
Subscribe to:
Posts (Atom)