From tgelder@phil.indiana.edu Fri Dec  1 02:11:35 1995
Received: from lucy.cs.wisc.edu by sea.cs.wisc.edu; Fri, 1 Dec 95 02:11:31 -0600; AA15317
Message-Id: <9512010811.AA12992@lucy.cs.wisc.edu>
Received: from TELNET-1.SRV.CS.CMU.EDU by lucy.cs.wisc.edu; Fri, 1 Dec 95 02:11:29 -0600
Received: from TELNET-1.SRV.CS.CMU.EDU by telnet-1.srv.cs.CMU.EDU id ab28865;
          29 Nov 95 12:00:34 EST
Received: from GS151.SP.CS.CMU.EDU by TELNET-1.SRV.CS.CMU.EDU id ab28863;
          29 Nov 95 11:40:10 EST
Received: from GS151.SP.CS.CMU.EDU by GS151.SP.CS.CMU.EDU id aa26506;
          29 Nov 95 11:40 EST
Received: from CS.CMU.EDU by B.GP.CS.CMU.EDU id aa24436; 29 Nov 95 0:18:50 EST
Received: from tarski.phil.indiana.edu by CS.CMU.EDU id aa05984;
          29 Nov 95 0:17:58 EST
Received: from [129.79.134.24] (tgelder.phil.indiana.edu) by phil.indiana.edu
	(4.1/9.7jsm) id AA16890; Wed, 29 Nov 95 00:17:54 EST
Mime-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Date: Wed, 29 Nov 1995 00:20:42 -0500
To: connectionists@cs.cmu.edu
From: Tim van Gelder <tgelder@phil.indiana.edu>
Subject:  'Mind as Motion' annct & web page

    Book announcement ::: Available now.

  `MIND AS MOTION: EXPLORATIONS IN THE DYNAMICS OF COGNITION'

        edited by Robert Port and Tim van Gelder
                Bradford Books/MIT Press.

>From the dust jacket: 

        `Mind as Motion is the first comprehensive presentation of the
dynamical approach to cognition. It contains a representative sampling of
original current research on topics such as perception, motor control,
speech and language, decision making, and development. Included are
chapters by pioneers of the approach, as well as others applying the tools
of dynamics to a wide range of new problems.  Throughout, particular
attention is paid to the philosophical foundations of this radical new
research program.
        Mind as Motion provides a conceptual and historical overview of the
dynamical approach, a tutorial introduction to dynamics for cognitive
scientists, and a glossary covering the most frequently used terms. Each
chapter includes an introduction by the editors, outlining its main ideas
and placing it in context, and a guide to further reading.'

           668 pages, 139 illustrations
           ISBN 0-262-16150-8
           $60.00 US

For further information including the full text of the preface and a sample
chapter introduction, see the web page at MIT Press:

   http://www-mitpress.mit.edu/mitp/recent-books/cog/mind-as-motion.html

_________________________________________________ 

   Chapter Titles and Authors

1.Tim van Gelder & Robert Port. 
        Introduction: Its About Time: An Overview of the Dynamical 
                Approach to Cognition.
2. Alec Norton.
         Dynamics: A Tutorial Introduction for Cognitive Scientists.
3. Esther Thelen.
        Time Scale Dynamics and the Development of an Embodied
                Cognition. 
4. Jerome Busemeyer & James Townsend
        Dynamic Representation of Decision Making
5. Randall Beer 
        Computational and Dynamical Languages for Autonomous Agents
6. Elliot Saltzman
        Dynamics and Coordinate Systems in Skilled Sensorimotor
                Activity
7. Catherine Browman & Louis Goldstein
        Dynamics and Articulatory Phonology
8. Jeffrey Elman
        Language as a Dynamical System
9. Jean Petitot
        Morphodynamics and Attractor Syntax
10. Jordan Pollack
        The Induction of Dynamical Recognizers
11. Paul van Geert 
        Growth Dynamics in Development
12. Robert Port, Fred Cummins & Devin McAuley
        Naive Time, Temporal Patterns and Human Audition
13. Michael Turvey & Claudia Carello
         Some Dynamical Themes in Perception and Action
14. Geoffrey Bingham
        Dynamics and the Problem of Event Recognition
15. Stephen Grossberg
        Neural Dynamics of Motion Perception, Recognition
                Learning, and Spatial Attention 
16. Mary Ann Metzger
        Multiprocess Models Applied to Cognitive and Behavioral Dynamics
17. Steven Reidbord & Dana Redington
        The Dynamics of Mind and Body During Clinical Interviews: 
                Current Trends, Potential, and Future Directions
18. Marco Giunti
        Dynamical Models of Cognition
19. Glossary of Terminology in Dynamics

_________________________________________________ 
At MIT Press, orders can be made by email at:
   mitpress-orders@mit.edu
For general information re MIT Press, see:
   http://www-mitpress.mit.edu

_________________________________________________
The editors welcome enquiries, discussion, critical feedback, etc.

Robert Port                     Tim van Gelder
Department of Linguistics       Department of Philosophy
Indiana University              University of Melbourne
Bloomington IN 47405            Parkville 3052 VIC
USA                             AUSTRALIA
port@indiana.edu                tgelder@ariel.unimelb.edu.au


From josh@faline.bellcore.com Fri Dec  1 02:11:37 1995
Received: from lucy.cs.wisc.edu by sea.cs.wisc.edu; Fri, 1 Dec 95 02:11:34 -0600; AA15321
Received: from TELNET-1.SRV.CS.CMU.EDU by lucy.cs.wisc.edu; Fri, 1 Dec 95 02:11:31 -0600
Received: from TELNET-1.SRV.CS.CMU.EDU by telnet-1.srv.cs.CMU.EDU id aa00751;
          30 Nov 95 13:48:53 EST
Received: from GS151.SP.CS.CMU.EDU by TELNET-1.SRV.CS.CMU.EDU id aa00749;
          30 Nov 95 13:27:42 EST
Received: from GS151.SP.CS.CMU.EDU by GS151.SP.CS.CMU.EDU id aa27157;
          30 Nov 95 13:27 EST
Received: from CS.CMU.EDU by B.GP.CS.CMU.EDU id aa04514; 29 Nov 95 13:20:35 EST
Received: from faline.bellcore.com by CS.CMU.EDU id aa10766;
          29 Nov 95 13:20:03 EST
Received: (from josh@localhost) by faline.bellcore.com (8.6.9/8.6.10) id NAA15771 for connectionists@cs.cmu.edu; Wed, 29 Nov 1995 13:19:47 -0500
Date: Wed, 29 Nov 1995 13:19:47 -0500
From: Joshua Alspector <josh@faline.bellcore.com>
Message-Id: <199511291819.NAA15771@faline.bellcore.com>
To: connectionists@cs.cmu.edu
Subject: Research associate in neuromorphic electronics


             RESEARCH ASSOCIATE IN NEUROMORPHIC ELECTRONICS

There is an anticipated position in the electrical and computer
engineering department at the University of Colorado at
Colorado Springs for a postdoctoral research associate in the area
of neural learning microchips.  The successful candidate will have
experience in analog and digital VLSI design and test and be comfortable
working at the system level in a UNIX/C/C++ environment.

The project will involve applying an existing VME-based
neural network learning system to several demanding problems in
signal processing.  These include adaptive non-linear equalization
of underwater acoustic communication channels and magnetic recording
channels.  It is likely also to involve integrating the learning electronics
with micro-machined sonic transducers directly on silicon.

Please send a curriculum vita, names and addresses of at least
three referees, and and copies of some representative publications to:

Prof. Joshua Alspector
Univ. of Colorado at Col. Springs
Dept. of Elec. & Comp. Eng.
P.O. Box 7150
Colorado Springs, CO 80933-7150
From ajit@austin.ibm.com Fri Dec  1 02:11:39 1995
Received: from lucy.cs.wisc.edu by sea.cs.wisc.edu; Fri, 1 Dec 95 02:11:35 -0600; AA15325
Received: from TELNET-1.SRV.CS.CMU.EDU by lucy.cs.wisc.edu; Fri, 1 Dec 95 02:11:33 -0600
Received: from TELNET-1.SRV.CS.CMU.EDU by telnet-1.srv.cs.CMU.EDU id ab00773;
          30 Nov 95 14:00:01 EST
Received: from GS151.SP.CS.CMU.EDU by TELNET-1.SRV.CS.CMU.EDU id aa00755;
          30 Nov 95 13:30:04 EST
Received: from GS151.SP.CS.CMU.EDU by GS151.SP.CS.CMU.EDU id aa27165;
          30 Nov 95 13:29 EST
Received: from CS.CMU.EDU by B.GP.CS.CMU.EDU id aa05916; 29 Nov 95 14:47:30 EST
Received: from netmail1.austin.ibm.com by CS.CMU.EDU id aa11877;
          29 Nov 95 14:46:31 EST
Received: from ding.austin.ibm.com (ding.austin.ibm.com [129.35.221.140]) by netmail1.austin.ibm.com (8.6.12/8.6.11) with SMTP id NAA52534 for <Connectionists@CS.CMU.EDU>; Wed, 29 Nov 1995 13:46:04 -0600
Received: by ding.austin.ibm.com (AIX 3.2/UCB 5.64/4.03-client-2.6)
          for Connectionists@CS.CMU.EDU at austin.ibm.com; id AA12676; Wed, 29 Nov 1995 13:46:04 -0600
Date: Wed, 29 Nov 1995 13:46:04 -0600
From: Dingankar <ajit@austin.ibm.com>
Message-Id: <9511291946.AA12676@ding.austin.ibm.com>
To: Connectionists@CS.cmu.edu
Subject: "Network Approximation of Dynamical Systems" - Neuroprose paper
Reply-To: ajit@austin.ibm.com
Organization: IBM, Austin
Work: (512) 838-6850
Fax: (512) 838-5882
Ftp-Host: archive.cis.ohio-state.edu
Ftp-Filename: /pub/neuroprose/dingankar.tensor-products.ps.Z


**DO NOT FORWARD TO OTHER GROUPS**
Sorry, no hardcopies available.
6 pages.

	Greetings!

	The following invited paper will be presented at NOLTA'95 next
month.  The compressed PostScript file is available in the Neuroprose
archive; the details (URL, bibtex entry and abstract) follow.

	Thanks,
	Ajit
------------------------------------------------------------------------------
URL:
ftp://archive.cis.ohio-state.edu/pub/neuroprose/dingankar.tensor-products.ps.Z

BiBTeX entry:
@INPROCEEDINGS{atd:nolta-95,
	AUTHOR		="Dingankar, Ajit T. and Sandberg, Irwin W.",
	TITLE		="{Network Approximation of Dynamical Systems}",
	BOOKTITLE	="Proceedings of the International Symposium
		 on Nonlinear Theory and its Applications (NOLTA'95)",
	YEAR		="1995",
	EDITOR		="",
	PAGES		="",
	ORGANIZATION	="",
	PUBLISHER	="",
	ADDRESS		="Las Vegas, Nevada",
	MONTH		="December 10--14"
}

		Network Approximation of Dynamical Systems
		------------------------------------------
				ABSTRACT
	We consider the problem of approximating any member of a large
class of input-output operators of time-varying nonlinear dynamical
systems.  We introduce a family of ``tensor product" dynamical neural
networks, and show that a certain continuity condition is necessary
and sufficient for the existence of arbitrarily good approximations
using this family.



------------------------------------------------------------------------------
Ajit T. Dingankar			|		ajit@austin.ibm.com 
IBM Corporation, Internal Zip 4359	|		Work: (512) 838-6850
11400 Burnet Road, Austin, TX 78758	|		Fax : (512) 838-5882

From krose@fwi.uva.nl Fri Dec  1 23:22:48 1995
Received: from lucy.cs.wisc.edu by sea.cs.wisc.edu; Fri, 1 Dec 95 23:22:44 -0600; AA02735
Received: from TELNET-1.SRV.CS.CMU.EDU by lucy.cs.wisc.edu; Fri, 1 Dec 95 23:22:42 -0600
Received: from TELNET-1.SRV.CS.CMU.EDU by telnet-1.srv.cs.CMU.EDU id aa00773;
          30 Nov 95 13:58:54 EST
Received: from GS151.SP.CS.CMU.EDU by TELNET-1.SRV.CS.CMU.EDU id aa00753;
          30 Nov 95 13:29:16 EST
Received: from GS151.SP.CS.CMU.EDU by GS151.SP.CS.CMU.EDU id aa27161;
          30 Nov 95 13:28 EST
Received: from CS.CMU.EDU by B.GP.CS.CMU.EDU id aa21440; 30 Nov 95 8:27:34 EST
Received: from mail.fwi.uva.nl by CS.CMU.EDU id aa18506; 30 Nov 95 8:26:16 EST
Received: from carol.fwi.uva.nl
          by mail.fwi.uva.nl with ESMTP (sendmail 8.6.12/config 5.15).
          id OAA14649; Thu, 30 Nov 1995 14:25:12 +0100
Return-Path: <krose@fwi.uva.nl>
Received: from localhost
          by carol.fwi.uva.nl (sendmail 8.6.12/config 5.15).
          id OAA08623; Thu, 30 Nov 1995 14:25:11 +0100
Message-Id: <199511301325.OAA08623@carol.fwi.uva.nl>
From: Ben Krose <krose@fwi.uva.nl>
X-Organisation: Faculty of Mathematics, Computer Science, Physics & Astronomy
                University of Amsterdam
                Kruislaan 403
                NL-1098 SJ Amsterdam
                The Netherlands
X-Phone:        +31 20 525 7463
X-Fax:          +31 20 525 7490
Subject: paper announcement
To: Connectionists@CS.cmu.edu
Date: Thu, 30 Nov 1995 14:25:10 +0100 (MET)
X-Mailer: ELM [version 2.4 PL23]
Mime-Version: 1.0
Content-Type: text/plain; charset=US-ASCII
Content-Transfer-Encoding: 7bit
Content-Length: 2014      


The following paper, published at the International Conference on Artificial
Neural Networks 1995 (ICANN*95), Springer Verlag, is available from:


http://www.fwi.uva.nl/research/neuro/publications/publications.html
or
ftp://ftp.fwi.uva.nl/pub/computer-systems/aut-sys/reports/VysGroKro95.ps.gz

The paper is 6 pages long, in compressed form 52kB.

Sorry, no hardcopies available.


===========================================================================

``Orthogonal incremental learning of a feedforward network'' 
V. Vysniauskas, F.C.A. Groen and B.J.A. Kr\"ose (1995) 
Proceedings of the 1995 International Conference on Neural Networks, 
Edited by Fogelman-Soulie and Gallinari, pp. 311-317. 

Abstract 

Orthogonal incremental learning (OIL) is a new approach of incremental 
training for a feedforward network with a single hidden layer. OIL is 
based on the idea to describe the output weights (but not the hidden 
nodes) as a set of orthogonal basis functions. Hidden nodes are treated 
just as the orthogonal representation of the network in the output weights 
domain. We showed that the network training can be performed incrementally, 
one node at time, and there is no need to use an additional constraint to 
support a consistent optimization among the hidden nodes. An advantage of 
OIL over existing algorithms is extremely fast learning. This approach
can be also easily extended to build-up incrementally an arbitrary function 
as a linear composition of adjustable functions which are not necessarily 
orthogonal. We tested this approach on a standard "two-spirals" benchmark 
problem to build incrementally a feedforward network with a single layer 
of Gaussian units. 


============================================================================

We welcome your comments,



Ben Kr\"ose
Faculty of Mathematics and Computer Science
University of Amsterdam
Kruislaan 403
1098 SJ Amsterdam
Netherlands
tel: +31 20 525 7520/7463

krose@fwi.uva.nl
http://www.fwi.uva.nl/research/neuro/


From dhw@santafe.edu Sat Dec  2 10:04:23 1995
Received: from lucy.cs.wisc.edu by sea.cs.wisc.edu; Sat, 2 Dec 95 10:04:19 -0600; AA08812
Received: from TELNET-1.SRV.CS.CMU.EDU by lucy.cs.wisc.edu; Sat, 2 Dec 95 10:04:16 -0600
Received: from TELNET-1.SRV.CS.CMU.EDU by telnet-1.srv.cs.CMU.EDU id aa02535;
          1 Dec 95 12:22:59 EST
Received: from GS151.SP.CS.CMU.EDU by TELNET-1.SRV.CS.CMU.EDU id aa02533;
          1 Dec 95 12:03:14 EST
Received: from GS151.SP.CS.CMU.EDU by GS151.SP.CS.CMU.EDU id aa27944;
          1 Dec 95 12:02 EST
Received: from CS.CMU.EDU by B.GP.CS.CMU.EDU id av13873; 1 Dec 95 11:27:53 EST
Received: from sfi.santafe.edu by CS.CMU.EDU id aa29741; 1 Dec 95 11:20:53 EST
Received: from santafe (santafe.santafe.edu) by sfi.santafe.edu (4.1/SMI-4.1)
	id AA27395; Fri, 1 Dec 95 09:18:19 MST
Date: Fri, 1 Dec 95 09:18:19 MST
From: David Wolpert <dhw@santafe.edu>
Message-Id: <9512011618.AA27395@sfi.santafe.edu>
To: Connectionists@CS.cmu.edu
Subject: Correcting misunderstandings about NFL


This posting is to correct some misunderstandings that were recently
posted concerning the NFL theorems. I also draw attention to some of
the incorrect interpretations commonly ascribed to certain COLT
results.

***


Joerg Lemm writes:

>>>
1.) If there is no relation between the function values
    on the test and training set
    (i.e. P(f(x_j)=y|Data) equal to the unconditional P(f(x_j)=y) ),
    then, having only training examples y_i = f(x_i) (=data) 
    from a given function, it is clear that I cannot learn anything 
    about values of the function at different arguments, 
    (i.e. for f(x_j), with x_j not equal to any x_i = nonoverlapping
test set).
>>>

Well put. Now here's the tough question: Vapnik *proves* that it is
unlikely (for large enough training sets and small enough VC dimension
generalizers) for error on the training set and full "generalization
error" to be grealy different. Regardless of the target. Using this,
Baum and Haussler even wrote a paper "What size net gives valid
generalization?" in which no assumptions whatsoever are made about the
target, and yet the authors are able to provide a response the
question of their title. HOW IS THAT POSSIBLE GIVEN WHAT YOU JUST
WROTE????

NFL is "obvious". And so are VC bounds on generalization error (well,
maybe not "obvious"). And so is the PAC "proof" of Occam's razor. And
yet the latter two bound generalization error (for those cases where
training set error is small enough) without making any assumptions
about the target. What gives?

The answer: The math of those works is correct. But far more care must
be exercised in the interpretation of that math than you will find in
those works. The care involves paying attention to what goes on the
right-hand side of the conditioning bars in one's probabilities, and
the implications of what goes there.  Unfortunately, such conditioning
bar are completely absent in those works...

(In fact, the sum-total of the difference between Bayesian and COLT
approaches to supervised batch learning lies in what's on the
right-hand side of those bars, but that's another story. See [2].)

As an example, it is widely realized that VC bounds suffer from being
worst-case.  However there is another hugely important caveat to those
bounds. The community as a whole simply is not aware of that caveat,
because the caveat concerns what goes on the right-hand side of the
conditioning bar, and this is NEVER made explicit.

This caveat is the fact that VC bounds do NOT concern

Pr(IID generalization error |
	observed error on the training set, training set size,
					VC dimension of the generalizer).

But you wouldn't know that to read the claims made on behalf of those
bounds ...

To give one simple example of the ramifications of this: Let's say you
have a favorite low-VC generalizer. And in the course of your career
you parse though learning problems, either explicitly or (far more
commonly) without even thinking about it. When you come across one
with a large training set on which your generalizer has small
generalization error, you want to invoke Vapnik to say you have
assuraces about full generalization error.

Well, sorry. You don't and you can't. You simply can't escape Bayes by
using confidence intervals. Confidence intervals in general (not just
in VC work) have the annoying property that as soon as you try to use
them, very often you contradict the underlying statistical assumptions
behind them. Details are in [1] and in the discussion of "We-Learn-It
Inc." in [2].



>>>
2.) We are considering two of those (influence) relations P(f(x_j)=y|Data):
    one, named A, for the true nature (=target) and one, named B, for our 
    model under study (=generalizer).
    Let P(A and B) be the joint probability distribution for the
    influence relations for target and generalizer.

3.) Of course, we do not know P(A and B), but in good old Bayesian tradition,
    we can construct a (hyper-)prior P(C) over the family of probability 
    distributions of the joint distributions C = P(A and B).
 
4.) NFL now uses the very special prior assumption
    P(A and B) = P(A)P(B)
>>>

If I understand you correctly, I would have to disagree. NFL also
holds with your P(C) being any prior assumption - more formally,
averaging over all priors, you get NFL. So the set of priors for which
your favorite algorithm does *worse than random* is just as large as
the set for which it does better. (In this sense, the uniform prior is
a typical prior, not a pathological one, out on the edge of the
space. It is certainly not a "very special prior".)

In fact, that's one of the major points of NFL - it's not to see what
life would be like if this or that were uniform, but to use such
uniformity as a mathematical tool, to get a handle on the underlying
geometry of inference, the size of the various spaces (e.g., the size
of the space of priors for which you lose to random), etc.

The math *starts* with NFL, and then goes on to many other things (see
[1]). It's only the beginning chapter of the text book.






>>>
I say that it is rational to believe 
(and David does so too, I think) that in real life cross-validation 
works better in more cases than  anti-cross-validation.
>>>

Oh, most definitely.

There are several issues here: 1) what gives with all the "prior-free"
general proofs of COLT, given NFL, 2) purely theoretical issues (e.g.,
as mentioned before, characterizing the relationship between target
and generalizers needed for xval. to beat anti-xval.) and 3) perhaps
most provocatively of all, seeing if NFL (and the associated
mathematical structure) can help you generalize in the real world
(e.g., with head-to-head minimax distinctions between generalizers).


***



Finally, Eric Baum weighs in:


>>>
Barak Pearlmutter remarked that saying
	We have *no* a priori reason to believe that targets with "low
	Kolmogorov complexity" (or anything else) are/not likely to
	occur in the real world.
(which I gather was a quote from David Wolpert?)
is akin to saying we have no a priori reason to believe there is non-random
structure in the world, which is not true, since we make great
predictions about the world.
>>>

Well, let's get a bit formal here. Take all the problems we've ever
tried to make "great predictions" on. Let's even say that these
problems were randomly chosen from those in the real world (i.e., no
selection effects of people simply not reporting when their
predictions were not so great). And let's for simplicity say that all
the predictions were generated by the same generalizer - the algorithm
in the brain of Eric Baum will do as a straw man.

Okay. Now take all those problems together and view them as one huge
training set. Better still, add in all the problems that Eric's
anscestors addressed, so that the success of his DNA is also taken
into account. That's still one training set. It's a huge one, but it's
tiny in comparison to the full spaces it lives in.

Saying we (Eric) makes "great predictions" simply means that the
xvalidation error of our generalizer (Eric) on that training set is
small. (You train on part of the data, and predict on the rest.)
Formally (!!!!!), this gives no assuraces whatsoever about any
behavior off-training-set. As I've stated before, without assumptions,
you cannot conclude that low xvalidation error leads to low
off-training set generalization error. And of course, each passing
second, each new scene you view, is "off-training-set".

The fallacy in Eric's claim was noted all the way back by
Hume. Success at inductive inference cannot formally establish the
utility of using inductive inference. To claim that it can you have to
invoke inductive inference, and that, as any second grader can tell
you, is circular reasoning.


Practically speaking of course, none of this is a concern in the real
world. We are all (me included) quite willing to conclude there is
structure in the real world.  But as was noted above, what we do in
practice is not the issue. The issue is one of theory.

***

It's very similar to high-energy physics. There are a bunch of
physical constants that, if only slightly varied, would (seem to) make
life impossible. Why do they have the values they have? Some invoke
the anthropic principle to answer this - we wouldn't be around if they
had other values. QED. But many find this a bit of a cop-out, and
search for something more fundamental. After all, you could have
stopped the progress of physics at any point in the past if you had
simply gotten everyone to buy into the anthropic principle at that
point in time.

Similarly with inductive inference. You could just cop out and say
"anthropic principle" - if inference were not possible, we wouldn't be
having this debate. But that's hardly a satisfying answer.


***


Eric goes on:

>>>
Consider the problem of learning
to predict the pressure of a gas from its temperature. Wolpert's theorem,
and his faith in our lack of prior about the world, predict,
that any learning algorithm whatever is as likely
to be good as any other. This is not correct.
>>>

To give two examples from just the past month, I'm sure MCI and
Coca-Cola would be astonished to know that the algorithms they're so
pleased with were designed for them by someone having "faith in our
lack of prior about the world".

Less glibly, let me address this claim about my "faith" with two
quotes from the NFL for supervised learning paper. The first is in the
introduction, and the second in a section entitled "On uniform
averaging". So neither is exactly hidden ...

1) "It cannot be emphasized enough that no claim is being made .. that
all algorithms are equivalent in the real world."

2) "The uniform sums over targets ... weren't chosen because there is
strong reason to believe that all targets are equally likely to arise
in practice. Indeed, in many respects it is absurd to ascribe such a
uniformity over possible targets to the real world. Rather the uniform
sums were chosen because such sums are a useful theoretical tool with
which to analyze supervised learning."


Finally, given that I'm mixing it up with Eric on NFL, I can't help
but quote the following from his "What size net gives valid
generalization" paper:

"We have given bounds (independent of the target) on the training set
size vs. neural net size need such that valid generalization can be
expected." 

(Parenthetical comment added - and true.)

Nowhere in the paper is there any discussion whatsoever of the
apparent contradiction between this statement and NFL-type concerns.
Indeed, as mentioned above, with only the conditioning-bar-free
mathematics in Eric's paper, there is no way to resolve the
contradiction. In this particular sense, that paper is extremely
misleading. (See discussion above on misinterpretations of Vapnik's
results.)



>>>>
Creatures evolving in this "play world" would exploit this structure and
understand their world in terms of it. There are other things they would
find hard to predict. In fact, it may be mathematically valid to say that
one could mathematically construct equally many functions on which
these creatures would fail to make good predictions. But so what?
So would their competition. This is not relevant to looking for
one's key, which is best done under the lamppost, where one has a
hope of finding it. In fact, it doesn't seem that the play world
creatures would care about all these other functions at all.
>>>

I'm not sure I quite follow this. In particular, the comment about the
"competition" seems to be wrong.

Let me just carry further Eric's metaphor though, and point out though
that it makes a hell of a lot more sense to pull out a flashlight and
explore into the surrounding territory for your key than it does to
spend all your time with your head down, banging into the lamppost. And
NFL is such a flashlight.




David Wolpert


[1] The current versions of the NFL for supervised learning papers,
nfl.ps.1.Z and nfl.ps.2.Z, at ftp.santafe.edu, in pub/dhw_ftp.

[2] "The Relationship between PAC, the Statistical Physics Framework,
the Bayesian Framework, and the VC Framework", in *The Mathematics of
Generalization*, D. Wolpert Ed., Addison-Wesley, 1995.






From marco@McCulloch.Ing.UniFI.IT Sat Dec  2 19:56:10 1995
Received: from lucy.cs.wisc.edu by sea.cs.wisc.edu; Sat, 2 Dec 95 19:56:07 -0600; AA13356
Received: from TELNET-1.SRV.CS.CMU.EDU by lucy.cs.wisc.edu; Sat, 2 Dec 95 19:56:05 -0600
Received: from TELNET-1.SRV.CS.CMU.EDU by telnet-1.srv.cs.CMU.EDU id aa04736;
          2 Dec 95 18:49:42 EST
Received: from GS151.SP.CS.CMU.EDU by TELNET-1.SRV.CS.CMU.EDU id aa04734;
          2 Dec 95 18:32:43 EST
Received: from GS151.SP.CS.CMU.EDU by GS151.SP.CS.CMU.EDU id aa29589;
          2 Dec 95 18:31 EST
Received: from RI.CMU.EDU by B.GP.CS.CMU.EDU id aa18886; 1 Dec 95 16:14:08 EST
Received: from cesit1.unifi.it by RI.CMU.EDU id aa25781; 1 Dec 95 16:12:54 EST
Received: from McCulloch.Ing.UniFI.IT by CESIT1.UNIFI.IT (PMDF V5.0-4 #3688)
 id <01HYAWOSUHY8000Q3Z@CESIT1.UNIFI.IT> for Connectionists@CS.CMU.EDU; Fri,
 01 Dec 1995 18:25:02 +0100 (MET)
Received: by McCulloch.Ing.UniFI.IT (5.x/SMI-SVR4) id AA09634; Fri,
 01 Dec 1995 18:21:43 +0100
Date: Fri, 01 Dec 1995 18:21:43 +0100
From: Marco Gori <marco@McCulloch.Ing.UniFI.IT>
Subject: Italian Neural Network Society
To: Connectionists@CS.cmu.edu
Message-Id: <9512011721.AA09634@McCulloch.Ing.UniFI.IT>
Organization: DSI - University of Florence (Italy)
Content-Transfer-Encoding: 7BIT
X-Sun-Charset: US-ASCII


==============================================================
This is to announce a new web page describing the aims and the
activities of the Italian Neural Network Society.  The page is
hosted at the DSI Web server of the Dipartimento di  Sistemi e 
Informatica, Universita' di Firenze) at the following address:

http://www-dsi.ing.unifi.it/neural/siren

-- marco gori. 
===============================================================

