提交 · 1092cb219774a82b1f16781aec7b8d4ec727c981 · openanolis / cloud-kernel

11 7月, 2007 36 次提交

[NETLINK]: attr: add nested compat attribute type · 1092cb21

由 Patrick McHardy 提交于 6月 25, 2007

Add a nested compat attribute type that can be used to convert
attributes that contain a structure to nested attributes in a
backwards compatible way.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

1092cb21

[SKBUFF]: Keep track of writable header len of headerless clones · 334a8132

由 Patrick McHardy 提交于 6月 25, 2007

Currently NAT (and others) that want to modify cloned skbs copy them,
even if in the vast majority of cases its not necessary because the
skb is a clone made by TCP and the portion NAT wants to modify is
actually writable because TCP release the header reference before
cloning.

The problem is that there is no clean way for NAT to find out how
long the writable header area is, so this patch introduces skb->hdr_len
to hold this length. When a headerless skb is cloned skb->hdr_len
is set to the current headroom, for regular clones it is copied from
the original. A new function skb_clone_writable(skb, len) returns
whether the skb is writable up to len bytes from skb->data. To avoid
enlarging the skb the mac_len field is reduced to 16 bit and the
new hdr_len field is put in the remaining 16 bit.

I've done a few rough benchmarks of NAT (not with this exact patch,
but a very similar one). As expected it saves huge amounts of system
time in case of sendfile, bringing it down to basically the same
amount as without NAT, with sendmsg it only helps on loopback,
probably because of the large MTU.

Transmit a 1GB file using sendfile/sendmsg over eth0/lo with and
without NAT:

- sendfile eth0, no NAT:	sys     0m0.388s
- sendfile eth0, NAT:		sys     0m1.835s
- sendfile eth0: NAT + path:	sys     0m0.370s	(~ -80%)

- sendfile lo, no NAT:		sys     0m0.258s
- sendfile lo, NAT:		sys     0m2.609s
- sendfile lo, NAT + patch:	sys     0m0.260s	(~ -90%)

- sendmsg eth0, no NAT:		sys     0m2.508s
- sendmsg eth0, NAT:		sys     0m2.539s
- sendmsg eth0, NAT + patch:	sys     0m2.445s	(no change)

- sendmsg lo, no NAT:		sys	0m2.151s
- sendmsg lo, NAT:		sys     0m3.557s
- sendmsg lo, NAT + patch:	sys     0m2.159s	(~ -40%)

I expect other users can see a similar performance improvement,
packet mangling iptables targets, ipip and ip_gre come to mind ..
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

334a8132

[NET]: qdisc_restart - couple of optimizations. · e50c41b5

由 Krishna Kumar 提交于 6月 24, 2007

Changes :

- netif_queue_stopped need not be called inside qdisc_restart as
  it has been called already in qdisc_run() before the first skb
  is sent, and in __qdisc_run() after each intermediate skb is
  sent (note : we are the only sender, so the queue cannot get
  stopped while the tx lock was got in the ~LLTX case).

- BUG_ON((int) q->q.qlen < 0) was a relic from old times when -1
  meant more packets are available, and __qdisc_run used to loop
  when qdisc_restart() returned -1. During those days, it was
  necessary to make sure that qlen is never less than zero, since
  __qdisc_run would get into an infinite loop if no packets are on
  the queue and this bug in qdisc was there (and worse - no more
  skbs could ever get queue'd as we hold the queue lock too). With
  Herbert's recent change to return values, this check is not
  required.  Hopefully Herbert can validate this change. If at all
  this is required, it should be added to skb_dequeue (in failure
  case), and not to qdisc_qlen.
Signed-off-by: NKrishna Kumar <krkumar2@in.ibm.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

e50c41b5

[NET]: qdisc_restart - readability changes plus one bug fix. · 6c1361a6

由 Krishna Kumar 提交于 6月 24, 2007

New changes :

- Incorporated Peter Waskiewicz's comments.
- Re-added back one warning message (on driver returning wrong value).

Previous changes :

- Converted to use switch/case code which looks neater.

- "if (ret == NETDEV_TX_LOCKED && lockless)" is buggy, and the lockless
  check should be removed, since driver will return NETDEV_TX_LOCKED only
  if lockless is true and driver has to do the locking. In the original
  code as well as the latest code, this code can result in a bug where
  if LLTX is not set for a driver (lockless == 0) but the driver is written
  wrongly to do a trylock (despite LLTX being set), the driver returns
  LOCKED. But since lockless is zero, the packet is requeue'd instead of
  calling collision code which will issue warning and free up the skb.
  Instead this skb will be retried with this driver next time, and the same
  result will ensue. Removing this check will catch these driver bugs instead
  of hiding the problem. I am keeping this change to readability section
  since :
  	a. it is confusing to check two things as it is; and
  	b. it is difficult to keep this check in the changed 'switch' code.

- Changed some names, like try_get_tx_pkt to dev_dequeue_skb (as that is
  the work being done and easier to understand) and do_dev_requeue to
  dev_requeue_skb, merged handle_dev_cpu_collision and tx_islocked to
  dev_handle_collision (handle_dev_cpu_collision is a small routine with only
  one caller, so there is no need to have two separate routines which also
  results in getting rid of two macros, etc.

- Removed an XXX comment as it should never fail (I suspect this was related
  to batch skb WIP, Jamal ?). Converted some functions to original coding
  style of having the return values and the function name on same line, eg
  prio2list.
Signed-off-by: NKrishna Kumar <krkumar2@in.ibm.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

6c1361a6

[CCID3]: Fix a bug in the send time processing · 49d66a70

由 Gerrit Renker 提交于 6月 16, 2007

ccid3_hc_tx_send_packet currently returns 0 when the time difference between
current time and t_nom is less than 1000 microseconds.

In this case the packet is sent immediately; but, unlike other packets that can
be emitted on first attempt, it will not have its window counter updated and
its options set as required. This is a bug.

Fix: Require the time difference to be at least 1000 microseconds. The
algorithm then converges: time differences > 1000 microseconds trigger the
timer in dccp_write_xmit; after timer expiry this function is tried again; when
the time difference is less than 1000, the packet will have its options added
and window counter updated as required.
Signed-off-by: NGerrit Renker <gerrit@erg.abdn.ac.uk>
Signed-off-by: NArnaldo Carvalho de Melo <acme@ghostprotocols.net>

49d66a70

[CCID3]: Sending time: update to ktime_t · 8132da4d

由 Gerrit Renker 提交于 6月 16, 2007

This updates the computation of t_nom and t_last_win_count to use the newer
gettimeofday interface.

Committer note: used ktime_to_timeval to set the 'now' variable to t_ld in
                ccid3hctx_no_feedback_timer
Signed-off-by: NGerrit Renker <gerrit@erg.abdn.ac.uk>
Signed-off-by: NArnaldo Carvalho de Melo <acme@ghostprotocols.net>

8132da4d

loss_interval: make struct dccp_li_hist_entry private · dd36a9ab

由 Arnaldo Carvalho de Melo 提交于 5月 28, 2007

net/dccp/ccids/lib/loss_interval.c is the only place where this struct is used.
Signed-off-by: NArnaldo Carvalho de Melo <acme@redhat.com>

dd36a9ab

loss_interval: Nuke dccp_li_hist · cc4d6a3a

由 Arnaldo Carvalho de Melo 提交于 5月 28, 2007

It had just a slab cache, so, for the sake of simplicity just make
dccp_trfc_lib module init routine create the slab cache, no need for users of
the lib to create a private loss_interval object.
Signed-off-by: NArnaldo Carvalho de Melo <acme@redhat.com>

cc4d6a3a

A
loss_interval: Make dccp_li_hist_entry_{new,delete} private · c70b729e
由 Arnaldo Carvalho de Melo 提交于 5月 28, 2007
```
Not used outside the loss_interval code anymore.
Signed-off-by: NArnaldo Carvalho de Melo <acme@redhat.com>
```
c70b729e
A
loss_interval: unexport dccp_li_hist_interval_new · 8c281780
由 Arnaldo Carvalho de Melo 提交于 5月 28, 2007
```
Now its only used inside the loss_interval code.
Signed-off-by: NArnaldo Carvalho de Melo <acme@redhat.com>
```
8c281780

[DCCP] loss_interval: Move ccid3_hc_rx_update_li to loss_interval · cc0a910b

由 Arnaldo Carvalho de Melo 提交于 6月 14, 2007

Renaming it to dccp_li_update_li.

Also based on previous work by Ian McDonald.
Signed-off-by: NArnaldo Carvalho de Melo <acme@ghostprotocols.net>

cc0a910b

[CCID3]: Pass ccid3_li_hist to ccid3_hc_rx_update_li · 878ac600

由 Arnaldo Carvalho de Melo 提交于 6月 14, 2007

Now ccid3_hc_rx_update_li is ready to be moved to
net/dccp/ccids/lib/loss_interval, it uses the same interface as the other
functions there.
Signed-off-by: NArnaldo Carvalho de Melo <acme@redhat.com>

878ac600

Remove accesses to ccid3_hc_rx_sock in ccid3_hc_rx_{update,calc_first}_li · d83258a3

由 Arnaldo Carvalho de Melo 提交于 5月 28, 2007

This is a preparatory patch for moving these loss interval functions from
net/dccp/ccids/ccid3.c to net/dccp/ccids/lib/loss_interval.c.

Based on a patch by Ian McDonald.
Signed-off-by: NArnaldo Carvalho de Melo <acme@ghostprotocols.net>

d83258a3

loss_interval: Fix timeval initialisation · 6bc7efe8

由 Ian McDonald 提交于 5月 28, 2007

When compiling with EXTRA_CFLAGS=-W noticed that tstamp is not initialised
correctly in dccp_li_calc_first_li.
Signed-off-by: NArnaldo Carvalho de Melo <acme@ghostprotocols.net>
Signed-off-by: NIan McDonald <ian.mcdonald@jandi.co.nz>

6bc7efe8

Fix dccp_sum_coverage · e961811f

由 Ian McDonald 提交于 5月 28, 2007

When compiling with EXTRA_CFLAGS=-W notice that we have signed/unsigned issue
in dccp.h.
Signed-off-by: NArnaldo Carvalho de Melo <acme@ghostprotocols.net>
Signed-off-by: NIan McDonald <ian.mcdonald@jandi.co.nz>

e961811f

ccid3: Update copyrights · b2f41ff4

由 Ian McDonald 提交于 5月 28, 2007

Signed-off-by: NIan McDonald <ian.mcdonald@jandi.co.nz>
Signed-off-by: NArnaldo Carvalho de Melo <acme@ghostprotocols.net>

b2f41ff4

[VLAN]: Use rtnl_link API · 07b5b17e

由 Patrick McHardy 提交于 6月 13, 2007

Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

07b5b17e

P
[VLAN]: Introduce symbolic constants for flag values · a4bf3af4
由 Patrick McHardy 提交于 6月 13, 2007
```
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
a4bf3af4

[VLAN]: Keep track of number of QoS mappings · b020cb48

由 Patrick McHardy 提交于 6月 13, 2007

Keep track of the number of configured ingress/egress QoS mappings to
avoid iteration while calculating the netlink attribute size.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

b020cb48

[VLAN]: Use 32 bit value for skb->priority mapping · 734423cf

由 Patrick McHardy 提交于 6月 13, 2007

skb->priority has only 32 bits and even VLAN uses 32 bit values in its API.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

734423cf

[VLAN]: Return proper error codes in register_vlan_device · 2ae0bf69

由 Patrick McHardy 提交于 6月 13, 2007

The returned device is unused, return proper error codes instead and avoid
having the ioctl handler guess the error.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

2ae0bf69

[VLAN]: Move device registation to seperate function · e89fe42c

由 Patrick McHardy 提交于 6月 13, 2007

Move device registration and configuration of the underlying device to a
seperate function.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

e89fe42c

[VLAN]: Split up device checks · c1d3ee99

由 Patrick McHardy 提交于 6月 13, 2007

Move the checks of the underlying device to a seperate function.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

c1d3ee99

[VLAN]: Move vlan_group allocation to seperate function · 42429aae

由 Patrick McHardy 提交于 6月 13, 2007

Move group allocation to a seperate function to clean up the code a bit
and allocate groups before registering the device. Device registration
is globally visible and causes netlink events, so we shouldn't fail
afterwards.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

42429aae

[VLAN]: Move some device intialization code to dev->init callback · 2f4284a4

由 Patrick McHardy 提交于 6月 13, 2007

Move some device initialization code to new dev->init callback to make
it shareable with netlink. Additionally this fixes a minor bug, dev->iflink
is set after registration, which causes an incorrect value in the initial
netlink message.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

2f4284a4

[VLAN]: Convert name-based configuration functions to struct netdevice * · c17d8874

由 Patrick McHardy 提交于 6月 13, 2007

Move the device lookup and checks to the ioctl handler under the RTNL and
change all name-based interfaces to take a struct net_device * instead.

This allows to use them from a netlink interface, which identifies devices
based on ifindex not name. It also avoids races between the ioctl interface
and the (upcoming) netlink interface since now all changes happen under the
RTNL.

As a nice side effect this greatly simplifies error handling in the helper
functions and fixes a number of incorrect error codes like -EINVAL for
device not found.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

c17d8874

[RTNETLINK]: Link creation API · 38f7b870

由 Patrick McHardy 提交于 6月 13, 2007

Add rtnetlink API for creating, changing and deleting software devices.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

38f7b870

[RTNETLINK]: Split up rtnl_setlink · 0157f60c

由 Patrick McHardy 提交于 6月 13, 2007

Split up rtnl_setlink into a function performing validation and a function
performing the actual changes. This allows to share the modifcation logic
with rtnl_newlink, which is introduced by the next patch.
Signed-off-by: NPatrick McHardy <kaber@trash.net>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

0157f60c

[MAC80211]: Add support for SIOCGIWRATE ioctl · b3d88ad4

由 Larry Finger 提交于 6月 10, 2007

At present, transmission rate information for mac80211 is available only
if verbose debugging is turned on, and then only in the logs. This patch
implements the SIOCGIWRATE ioctl, which adds the current transmission rate to
the output of iwconfig.
Signed-off-by: NLarry Finger <Larry.Finger@lwfinger.net>
Signed-off-by: NJohn W. Linville <linville@tuxdriver.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

b3d88ad4

[TCPv4]: Improve BH latency in /proc/net/tcp · a7ab4b50

由 Herbert Xu 提交于 6月 10, 2007

Currently the code for /proc/net/tcp disable BH while iterating
over the entire established hash table.  Even though we call
cond_resched_softirq for each entry, we still won't process
softirq's as regularly as we would otherwise do which results
in poor performance when the system is loaded near capacity.

This anomaly comes from the 2.4 code where this was all in a
single function and the local_bh_disable might have made sense
as a small optimisation.

The cost of each local_bh_disable is so small when compared
against the increased latency in keeping it disabled over a
large but mostly empty TCP established hash table that we
should just move it to the individual read_lock/read_unlock
calls as we do in inet_diag.
Signed-off-by: NHerbert Xu <herbert@gondor.apana.org.au>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

a7ab4b50

[NET_SCHED]: Cleanup readability of qdisc restart · c716a81a

由 Jamal Hadi Salim 提交于 6月 10, 2007

Over the years this code has gotten hairier. Resulting in many long
discussions over long summer days and patches that get it wrong.
This patch helps tame that code so normal people will understand it.

Thanks to Thomas Graf, Peter J. waskiewicz Jr, and Patrick McHardy
for their valuable reviews.
Signed-off-by: NJamal Hadi Salim <hadi@cyberus.ca>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

c716a81a

[TIPC]: Optimize stream send routine to avoid fragmentation · 05646c91

由 Allan Stephens 提交于 6月 10, 2007

This patch enhances TIPC's stream socket send routine so that
it avoids transmitting data in chunks that require fragmentation
and reassembly, thereby improving performance at both the
sending and receiving ends of the connection.

The "maximum packet size" hint that records MTU info allows
the socket to decide how big a chunk it should send; in the
event that the hint has become stale, fragmentation may still
occur, but the data will be passed correctly and the hint will
be updated in time for the following send.  Note: The 66060 byte
pseudo-MTU used for intra-node connections requires the send
routine to perform an additional check to ensure it does not
exceed TIPC"s limit of 66000 bytes of user data per chunk.
Signed-off-by: NAllan Stephens <allan.stephens@windriver.com>
Signed-off-by: NJon Paul Maloy <jon.maloy@ericsson.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

05646c91

[TIPC]: Use standard socket "not implemented" routines · 5eee6a6d

由 Allan Stephens 提交于 6月 10, 2007

This patch modifies TIPC's socket API to utilize existing
generic routines to indicate unsupported operations, rather
than adding similar TIPC-specific routines.
Signed-off-by: NAllan Stephens <allan.stephens@windriver.com>
Signed-off-by: NJon Paul Maloy <jon.maloy@ericsson.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

5eee6a6d

[TIPC]: Improved support for Ethernet traffic filtering · f3ec75f6

由 Allan Stephens 提交于 6月 10, 2007

This patch simplifies TIPC's Ethernet receive routine to take
advantage of information already present in each incoming sk_buff
indicating whether the packet was explicitly sent to the interface,
has been broadcast to all interfaces, or was picked up because the
interface is in promiscous mode.

This new approach also fixes the problem of TIPC accepting unwanted
traffic through UML's multicast-based Ethernet interfaces (which
deliver traffic in a promiscuous manner even if the interface is
not configured to be promiscuous).
Signed-off-by: NAllan Stephens <allan.stephens@windriver.com>
Signed-off-by: NJon Paul Maloy <jon.maloy@ericsson.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

f3ec75f6

D
[IPV4]: The scheduled removal of multipath cached routing support. · e06e7c61
由 David S. Miller 提交于 6月 10, 2007
```
With help from Chris Wedgwood.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
e06e7c61

bonding / ipv6: no addrconf for slaves separately from master · c2edacf8

由 Jay Vosburgh 提交于 7月 09, 2007

	At present, when a device is enslaved to bonding, if ipv6 is
active then addrconf will be initated on the slave (because it is closed
then opened during the enslavement processing).  This causes DAD and RS
packets to be sent from the slave.  These packets in turn can confuse
switches that perform ipv6 snooping, causing them to incorrectly update
their forwarding tables (if, e.g., the slave being added is an inactve
backup that won't be used right away) and direct traffic away from the
active slave to a backup slave (where the incoming packets will be
dropped).

	This patch alters the behavior so that addrconf will only run on
the master device itself.  I believe this is logically correct, as it
prevents slaves from having an IPv6 identity independent from the
master.  This is consistent with the IPv4 behavior for bonding.

	This is accomplished by (a) having bonding set IFF_SLAVE sooner
in the enslavement processing than currently occurs (before open, not
after), and (b) having ipv6 addrconf ignore UP and CHANGE events on
slave devices.

	The eql driver also uses the IFF_SLAVE flag.  I inspected eql,
and I believe this change is reasonable for its usage of IFF_SLAVE, but
I did not test it.
Signed-off-by: NJay Vosburgh <fubar@us.ibm.com>
Signed-off-by: NJeff Garzik <jeff@garzik.org>

c2edacf8

10 7月, 2007 1 次提交
- J
  sendfile: convert nfsd to splice_direct_to_actor() · cf8208d0
  由 Jens Axboe 提交于 6月 12, 2007
```
Signed-off-by: NJens Axboe <jens.axboe@oracle.com>
```
  cf8208d0
09 7月, 2007 1 次提交

[PATCH] softmac: use list_for_each_entry · 67c4f7aa

由 Akinobu Mita 提交于 5月 27, 2007

Cleanup using list_for_each_entry.

Cc: Johannes Berg <johannes@sipsolutions.net>
Cc: Joe Jezak <josejx@gentoo.org>
Cc: Daniel Drake <dsd@gentoo.org>
Signed-off-by: NAkinobu Mita <akinobu.mita@gmail.com>
Signed-off-by: NJohn W. Linville <linville@tuxdriver.com>

67c4f7aa

08 7月, 2007 1 次提交

Fix use-after-free oops in Bluetooth HID. · 1c39858b

由 David Woodhouse 提交于 7月 07, 2007

When cleaning up HIDP sessions, we currently close the ACL connection
before deregistering the input device. Closing the ACL connection
schedules a workqueue to remove the associated objects from sysfs, but
the input device still refers to them -- and if the workqueue happens to
run before the input device removal, the kernel will oops when trying to
look up PHYSDEVPATH for the removed input device.

Fix this by deregistering the input device before closing the
connections.
Signed-off-by: NDavid Woodhouse <dwmw2@infradead.org>
Acked-by: NMarcel Holtmann <marcel@holtmann.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

1c39858b

06 7月, 2007 1 次提交

[NETPOLL]: Fixups for 'fix soft lockup when removing module' · 25442caf

由 Jarek Poplawski 提交于 7月 05, 2007

>From my recent patch:

> >    #1
> >    Until kernel ver. 2.6.21 (including) cancel_rearming_delayed_work()
> >    required a work function should always (unconditionally) rearm with
> >    delay > 0 - otherwise it would endlessly loop. This patch replaces
> >    this function with cancel_delayed_work(). Later kernel versions don't
> >    require this, so here it's only for uniformity.

But Oleg Nesterov <oleg@tv-sign.ru> found:

> But 2.6.22 doesn't need this change, why it was merged?
> 
> In fact, I suspect this change adds a race,
...

His description was right (thanks), so this patch reverts #1.
Signed-off-by: NJarek Poplawski <jarkao2@o2.pl>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

25442caf

openanolis / cloud-kernel 大约 1 年 前同步成功

openanolis / cloud-kernel
大约 1 年前同步成功