Digging into the future of data mining

Crystal ball gazing


Comment The first thing to appreciate about data mining is that it should be thought of as R&D. That is, you do a bunch of research, some of which (but by no means all) is then deployable in the business. Moreover, some of it becomes so well established that it becomes a mass market product. For example, market basket analysis (which products have relationships to others) was once regarded as being as esoteric as anything else in data mining but is now so mainstream that it is embedded in all sorts of other environments. This is a trend that will continue, with techniques moving out of data mining R&D and into conventional deployment.

Historically, this move out of data mining has been to call centres, CRM, fraud and other standard applications. However, as complex event processing (CEP) engines take greater market share then we are likely to see increasing synergy with data mining. After all, CEP is essentially about identifying patterns and then detecting anomalies, which is exactly what data mining does.

There is a lot of hype about predictive analytics as opposed to data mining. If we take the case of market basket analysis, this is essentially saying that once we have identified that the sale of nappies is associated with beer sales (even if that is an urban myth) then we can make predictions about one based on the other. Useful, and certainly an increasing focus, but not really significantly different from what data mining has always been about.

Of course, there is also a trend to make data mining easier (and less costly) to do, but that is hardly surprising: it is common across the whole IT sector.

In my view, perhaps the most important trend is towards the integration of text mining and data mining. As yet, this is a relatively immature market but the fact is that most information held within business today is in unstructured format. While most of the discussion has been about Search that is simply about finding things related to a particular topic, while text mining is about finding patterns of information within text which, in the right context, is much more valuable. Moreover, with the advent of DB2 Viper we are likely to see the increased use of applications that employ both relational and XML-based information, in which case a combination of data and text mining makes sense.

While SPSS is one of the two major players in the data mining market it is the clear leader in the text mining space, not least because it is the dominant provider of market research software, and doing text mining on the back of the results of market research makes obvious sense. However, it is probably also the leading provider of combined text and data mining outside of this environment as well, so if I am right about the future of data mining, and its increased use with text capabilities, then SPSS is very well-placed.

SPSS is also in a good position because IBM has withdrawn the client component of its Intelligent Miner product and users thereof will be looking for a replacement offering, and SPSS has a much closer relationship with IBM than its major competitors, which it is looking to capitalise upon by picking up these users.

Copyright © 2006, IT-Analysis.com


Other stories you might like

  • It's 2022 and there are still malware-laden PDFs in emails exploiting bugs from 2017
    Crafty file names, encrypted malicious code, Office flaws – ah, it's like the Before Times

    HP's cybersecurity folks have uncovered an email campaign that ticks all the boxes: messages with a PDF attached that embeds a Word document that upon opening infects the victim's Windows PC with malware by exploiting a four-year-old code-execution vulnerability in Microsoft Office.

    Booby-trapping a PDF with a malicious Word document goes against the norm of the past 10 years, according to the HP Wolf Security researchers. For a decade, miscreants have preferred Office file formats, such as Word and Excel, to deliver malicious code rather than PDFs, as users are more used to getting and opening .docx and .xlsx files. About 45 percent of malware stopped by HP's threat intelligence team in the first quarter of the year leveraged Office formats.

    "The reasons are clear: users are familiar with these file types, the applications used to open them are ubiquitous, and they are suited to social engineering lures," Patrick Schläpfer, malware analyst at HP, explained in a write-up, adding that in this latest campaign, "the malware arrived in a PDF document – a format attackers less commonly use to infect PCs."

    Continue reading
  • New audio server Pipewire coming to next version of Ubuntu
    What does that mean? Better latency and a replacement for PulseAudio

    The next release of Ubuntu, version 22.10 and codenamed Kinetic Kudu, will switch audio servers to the relatively new PipeWire.

    Don't panic. As J M Barrie said: "All of this has happened before, and it will all happen again." Fedora switched to PipeWire in version 34, over a year ago now. Users who aren't pro-level creators or editors of sound and music on Ubuntu may not notice the planned change.

    Currently, most editions of Ubuntu use the PulseAudio server, which it adopted in version 8.04 Hardy Heron, the company's second LTS release. (The Ubuntu Studio edition uses JACK instead.) Fedora 8 also switched to PulseAudio. Before PulseAudio became the standard, many distros used ESD, the Enlightened Sound Daemon, which came out of the Enlightenment project, best known for its desktop.

    Continue reading
  • VMware claims 'bare-metal' performance on virtualized GPUs
    Is... is that why Broadcom wants to buy it?

    The future of high-performance computing will be virtualized, VMware's Uday Kurkure has told The Register.

    Kurkure, the lead engineer for VMware's performance engineering team, has spent the past five years working on ways to virtualize machine-learning workloads running on accelerators. Earlier this month his team reported "near or better than bare-metal performance" for Bidirectional Encoder Representations from Transformers (BERT) and Mask R-CNN — two popular machine-learning workloads — running on virtualized GPUs (vGPU) connected using Nvidia's NVLink interconnect.

    NVLink enables compute and memory resources to be shared across up to four GPUs over a high-bandwidth mesh fabric operating at 6.25GB/s per lane compared to PCIe 4.0's 2.5GB/s. The interconnect enabled Kurkure's team to pool 160GB of GPU memory from the Dell PowerEdge system's four 40GB Nvidia A100 SXM GPUs.

    Continue reading
  • Nvidia promises annual updates across CPU, GPU, and DPU lines
    Arm one year, x86 the next, and always faster than a certain chip shop that still can't ship even one standalone GPU

    Computex Nvidia's push deeper into enterprise computing will see its practice of introducing a new GPU architecture every two years brought to its CPUs and data processing units (DPUs, aka SmartNICs).

    Speaking on the company's pre-recorded keynote released to coincide with the Computex exhibition in Taiwan this week, senior vice president for hardware engineering Brian Kelleher spoke of the company's "reputation for unmatched execution on silicon." That's language that needs to be considered in the context of Intel, an Nvidia rival, again delaying a planned entry to the discrete GPU market.

    "We will extend our execution excellence and give each of our chip architectures a two-year rhythm," Kelleher added.

    Continue reading
  • Amazon puts 'creepy' AI cameras in UK delivery vans
    Big Bezos is watching you

    Amazon is reportedly installing AI-powered cameras in delivery vans to keep tabs on its drivers in the UK.

    The technology was first deployed, with numerous errors that reportedly denied drivers' bonuses after malfunctions, in the US. Last year, the internet giant produced a corporate video detailing how the cameras monitor drivers' driving behavior for safety reasons. The same system is now apparently being rolled out to vehicles in the UK. 

    Multiple camera lenses are placed under the front mirror. One is directed at the person behind the wheel, one is facing the road, and two are located on either side to provide a wider view. The cameras are monitored by software built by Netradyne, a computer-vision startup focused on driver safety. This code uses machine-learning algorithms to figure out what's going on in and around the vehicle.

    Continue reading

Biting the hand that feeds IT © 1998–2022