Recently, I have been talking with some software folks who have created parsers for auto-classification. These engines will take a text or website, scan the contents and see how it matches to a given taxonomy. These parsers use a variety of techniques including statistics, semantics, word location to see how a given document matches a taxonomy. These are very impressive, and I hope to use one of these for our taxonomy to classify our content. One of the unique traits of our taxonomy is that is multidimensional, and we use those dimensions to show the different ways that businesses operate in our economy. We have been able to leverage this to find buyers, sellers and comparables for Mergers and Acquisitions. One of the problems with this approach is that users who are searching our database need to have a fairly in depth knowledge of the taxonomy in order to find what they need. To solve this problem, we decided we needed to create meta-terms or coined phrases that would represent different search criteria to apply to our taxonomy. Instead of, going from text to taxonomy we are going from taxonomy to text. We also use all the synonyms for all the nodes in our taxonomy. So instead of having users select their criteria from a series of drop-down boxes, they can type in a Google-like text box and the software will auto-fill matches against the table of coined phrases generated from our taxonomy.
I have a live sample on our MandAsoft site here: http://mandasoft.com/segments/searchlob.aspx. A good example is accountant software which is a coined phrase that will search for software for accountants. Here are the results for that search http://mandasoft.com/segmentview.aspx?SearchID=LOB99.
Tell me what you think.
Follow the exploration on how to build a better taxonomy for business. I believe current taxonomies are currently lacking. Using a spatial concept of taxonomy, our team has built a taxonomy that is used to power research into Comparables for Mergers & Acquisitions in the evolving sector of Information, Software and Media, and I would like to share our thoughts on how to build a better taxonomy.
Showing posts with label Auto-Classification. Show all posts
Showing posts with label Auto-Classification. Show all posts
Wednesday, January 25, 2012
Tuesday, January 17, 2012
Taxonomy Evolution Conundrum
Our team has been developing our taxonomy for almost ten years now. Our goal is to classify businesses by looking at how they operate, who they serve, and what they do, and our focus has been on media and software businesses. Needless to say over the last ten years, there have been major changes to the media and software industries with the introduction of smart phones, tablet computers, cloud computing, SaaS, virtualization, etc. To handle this evolution of the content we are classifying, we need to make sure our framework was solid and that the taxonomy could change with abilities to add nodes, merge nodes, link nodes, and to make sure our classifications migrated with the changes. However, change is never apparent when it happens. When we saw the first business operating in Social Networking, we originally had them classified basically as forums of user generated content, as opposed to editorial content. But as the business and technology took off, and showed itself to be a new business model, we realized we had to add the term Social Networking to our taxonomy. Now our problem was that we had to go back and re-evaluate our companies that were classified as forums and see if they were really Social Networks. One way to fix this problem is to have an auto-classifier, and you set up a new set of rules to recognize Social Networking. Then you re-run the auto-classifier on those companies. But here is the conundrum, we noticed this evolution in business models because we had human eyes seeing the trend. How can you expect an auto-classifier to see that? What are your thoughts on this problem?
Wednesday, December 21, 2011
Challenges of Classifying a Business
One of the biggest challenges our team has building a taxonomy of businesses is the actual classifying of a business. The process at best is semi-automated. Our classifiers generally have to look at the website of a given business and try to determine what they do. The website, generally speaking, is not written to describe how a business operates, but rather to sell the business's products and services. Each website has its own content, and we used to concentrate on the "About Us" page where the business defines itself. Unfortunately, the "About Us" page is usually some sort of vague mission statement. We even looked at our own website, and found the wording written by our Marketing guy was so general that you would not know what we did! So our classifiers have learned to scan around many pages of a businesses site to learn what they do, how they do it and who they sell to. We have built some scrapers, but our results have been mediocre. We are now currently looking at ways to scrape company websites and intelligently gather info data to further automate the classification. The big problem we saw was that matching words on the website against words in our thesaurus gave us too many false positives. We are now looking at weighting the value of words in our thesaurus and how they matched in the past with verified classifications. Anyone explore these types of auto-classification.
Subscribe to:
Posts (Atom)