提交 · 94206125c4aac32e43c25bfe1b827e7ab993b7dc · openeuler / raspberrypi-kernel

12 7月, 2012 2 次提交

ipv4: Rearrange arguments to ip_rt_redirect() · 94206125

由 David S. Miller 提交于 7月 11, 2012

Pass in the SKB rather than just the IP addresses, so that policy
and other aspects can reside in ip_rt_redirect() rather then
icmp_redirect().
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

94206125

tcp: TCP Small Queues · 46d3ceab

由 Eric Dumazet 提交于 7月 11, 2012

This introduce TSQ (TCP Small Queues)

TSQ goal is to reduce number of TCP packets in xmit queues (qdisc &
device queues), to reduce RTT and cwnd bias, part of the bufferbloat
problem.

sk->sk_wmem_alloc not allowed to grow above a given limit,
allowing no more than ~128KB [1] per tcp socket in qdisc/dev layers at a
given time.

TSO packets are sized/capped to half the limit, so that we have two
TSO packets in flight, allowing better bandwidth use.

As a side effect, setting the limit to 40000 automatically reduces the
standard gso max limit (65536) to 40000/2 : It can help to reduce
latencies of high prio packets, having smaller TSO packets.

This means we divert sock_wfree() to a tcp_wfree() handler, to
queue/send following frames when skb_orphan() [2] is called for the
already queued skbs.

Results on my dev machines (tg3/ixgbe nics) are really impressive,
using standard pfifo_fast, and with or without TSO/GSO.

Without reduction of nominal bandwidth, we have reduction of buffering
per bulk sender :
< 1ms on Gbit (instead of 50ms with TSO)
< 8ms on 100Mbit (instead of 132 ms)

I no longer have 4 MBytes backlogged in qdisc by a single netperf
session, and both side socket autotuning no longer use 4 Mbytes.

As skb destructor cannot restart xmit itself ( as qdisc lock might be
taken at this point ), we delegate the work to a tasklet. We use one
tasklest per cpu for performance reasons.

If tasklet finds a socket owned by the user, it sets TSQ_OWNED flag.
This flag is tested in a new protocol method called from release_sock(),
to eventually send new segments.

[1] New /proc/sys/net/ipv4/tcp_limit_output_bytes tunable
[2] skb_orphan() is usually called at TX completion time,
  but some drivers call it in their start_xmit() handler.
  These drivers should at least use BQL, or else a single TCP
  session can still fill the whole NIC TX ring, since TSQ will
  have no effect.
Signed-off-by: NEric Dumazet <edumazet@google.com>
Cc: Dave Taht <dave.taht@bufferbloat.net>
Cc: Tom Herbert <therbert@google.com>
Cc: Matt Mathis <mattmathis@google.com>
Cc: Yuchung Cheng <ycheng@google.com>
Cc: Nandita Dukkipati <nanditad@google.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

46d3ceab

11 7月, 2012 15 次提交

ipv6: Move ipv6 twsk accessors outside of CONFIG_IPV6 ifdefs. · 48ee3569

由 David S. Miller 提交于 7月 11, 2012

Fixes build when ipv6 is disabled.
Reported-by: NFengguang Wu <wfg@linux.intel.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

48ee3569

ipv6: optimize ipv6 addresses compares · 1a203cb3

由 Eric Dumazet 提交于 7月 10, 2012

On 64 bit arches having efficient unaligned accesses (eg x86_64) we can
use long words to reduce number of instructions for free.

Joe Perches suggested to change ipv6_masked_addr_cmp() to return a bool
instead of 'int', to make sure ipv6_masked_addr_cmp() cannot be used
in a sorting function.
Signed-off-by: NEric Dumazet <edumazet@google.com>
Cc: Joe Perches <joe@perches.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

1a203cb3

D
ipv4: Remove inetpeer from routes. · f185071d
由 David S. Miller 提交于 7月 10, 2012
```
No longer used.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
f185071d

ipv4: Maintain redirect and PMTU info in struct rtable again. · 5943634f

由 David S. Miller 提交于 7月 10, 2012

Maintaining this in the inetpeer entries was not the right way to do
this at all.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

5943634f

D
rtnetlink: Remove ts/tsage args to rtnl_put_cacheinfo(). · 87a50699
由 David S. Miller 提交于 7月 10, 2012
```
Nobody provides non-zero values any longer.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
87a50699

inet: Kill FLOWI_FLAG_PRECOW_METRICS. · 3e12939a

由 David S. Miller 提交于 7月 10, 2012

No longer needed.  TCP writes metrics, but now in it's own special
cache that does not dirty the route metrics.  Therefore there is no
longer any reason to pre-cow metrics in this way.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

3e12939a

D
inet: Remove ->get_peer() method. · 16d18399
由 David S. Miller 提交于 7月 10, 2012
```
No longer used.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
16d18399
D
tcp: Remove tw->tw_peer · b6242b9b
由 David S. Miller 提交于 7月 10, 2012
```
No longer used.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
b6242b9b
D
tcp: Move timestamps from inetpeer to metrics cache. · 81166dd6
由 David S. Miller 提交于 7月 10, 2012
```
With help from Lin Ming.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
81166dd6
D
net: Kill set_dst_metric_rtt(). · 94334d5e
由 David S. Miller 提交于 7月 10, 2012
```
No longer used.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
94334d5e

tcp: Maintain dynamic metrics in local cache. · 51c5d0c4

由 David S. Miller 提交于 7月 10, 2012

Maintain a local hash table of TCP dynamic metrics blobs.

Computed TCP metrics are no longer maintained in the route metrics.

The table uses RCU and an extremely simple hash so that it has low
latency and low overhead.  A simple hash is legitimate because we only
make metrics blobs for fully established connections.

Some tweaking of the default hash table sizes, metric timeouts, and
the hash chain length limit certainly could use some tweaking.  But
the basic design seems sound.

With help from Eric Dumazet and Joe Perches.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

51c5d0c4

D
tcp: Abstract back handling peer aliveness test into helper function. · ab92bb2f
由 David S. Miller 提交于 7月 09, 2012
```
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
ab92bb2f
D
tcp: Move dynamnic metrics handling into seperate file. · 4aabd8ef
由 David S. Miller 提交于 7月 09, 2012
```
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
4aabd8ef

etherdevice: introduce eth_broadcast_addr · ad7eee98

由 Johannes Berg 提交于 7月 10, 2012

A lot of code has either the memset or an inefficient copy
from a static array that contains the all-ones broadcast
address. Introduce eth_broadcast_addr() to fill an address
with all ones, making the code clearer and allowing us to
get rid of some constant arrays.
Signed-off-by: NJohannes Berg <johannes.berg@intel.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

ad7eee98

ipv4: Fix crashes in fib_rules_tclass(). · e044a651

由 David S. Miller 提交于 7月 10, 2012

All paths assume, when CONFIG_IP_MULTIPLE_TABLES is enabled, that any
successful call to fib_lookup() will initialize the fib_result->r
value to something.

We violated that expectation in the new fib_lookup() fast path.
Reported-by: NOr Gerlitz <ogerlitz@mellanox.com>
Tested-by: NEric Dumazet <eric.dumazet@gmail.com>
Tested-by: NGreg Rose <gregory.v.rose@intel.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

e044a651

09 7月, 2012 2 次提交

netfilter: nf_ct_ecache: fix crash with multiple containers, one shutting down · 6bd0405b

由 Pablo Neira Ayuso 提交于 7月 05, 2012

Hans reports that he's still hitting:

BUG: unable to handle kernel NULL pointer dereference at 000000000000027c
IP: [<ffffffff813615db>] netlink_has_listeners+0xb/0x60
PGD 0
Oops: 0000 [#3] PREEMPT SMP
CPU 0

It happens when adding a number of containers with do:

nfct_query(h, NFCT_Q_CREATE, ct);

and most likely one namespace shuts down.

this problem was supposed to be fixed by:
70e9942f netfilter: nf_conntrack: make event callback registration per-netns

Still, it was missing one rcu_access_pointer to check if the callback
is set or not.
Reported-by: NHans Schillstrom <hans@schillstrom.com>
Signed-off-by: NPablo Neira Ayuso <pablo@netfilter.org>

6bd0405b

phylib: Support registering a bunch of drivers · d5bf9071

由 Christian Hohnstaedt 提交于 7月 04, 2012

If registering of one of them fails, all already registered drivers
of this module will be unregistered.

Use the new register/unregister functions in all drivers
registering more than one driver.

amd.c, realtek.c: Simplify: directly return registration result.

Tested with broadcom.c
All others compile-tested.
Signed-off-by: NChristian Hohnstaedt <chohnstaedt@innominate.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

d5bf9071

08 7月, 2012 4 次提交

net/mlx4: Implement promiscuous mode with device managed flow-steering · 592e49dd

由 Hadar Hen Zion 提交于 7月 05, 2012

The device managed flow steering API has three promiscuous modes:

1. Uplink - captures all the packets that arrive to the port.
2. Allmulti - captures all multicast packets arriving to the port.
3. Function port - for future use, this mode is not implemented yet.

Use these modes with the flow_attach and flow_detach firmware commands
according to the promiscuous state of the netdevice.
Signed-off-by: NHadar Hen Zion <hadarh@mellanox.co.il>
Signed-off-by: NOr Gerlitz <ogerlitz@mellanox.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

592e49dd

{NET, IB}/mlx4: Add device managed flow steering firmware API · 0ff1fb65

由 Hadar Hen Zion 提交于 7月 05, 2012

The driver is modified to support three operation modes.

If supported by firmware use the device managed flow steering
API, that which we call device managed steering mode. Else, if
the firmware supports the B0 steering mode use it, and finally,
if none of the above, use the A0 steering mode.

When the steering mode is device managed, the code is modified
such that L2 based rules set by the mlx4_en driver for Ethernet
unicast and multicast, and the IB stack multicast attach calls
done through the mlx4_ib driver are all routed to use the device
managed API.

When attaching rule using device managed flow steering API,
the firmware returns a 64 bit registration id, which is to be
provided during detach.

Currently the firmware is always programmed during HCA initialization
to use standard L2 hashing. Future work should be done to allow
configuring the flow-steering hash function with common, non
proprietary means.
Signed-off-by: NHadar Hen Zion <hadarh@mellanox.co.il>
Signed-off-by: NOr Gerlitz <ogerlitz@mellanox.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

0ff1fb65

net/mlx4_core: Add firmware commands to support device managed flow steering · 8fcfb4db

由 Hadar Hen Zion 提交于 7月 05, 2012

Add support for firmware commands to attach/detach a new device managed
steering mode. Such network steering rules allow the user to provide an
L2/L3/L4 flow specification to the firmware and have the device to steer
traffic that matches that specification to the provided QP.
Signed-off-by: NHadar Hen Zion <hadarh@mellanox.co.il>
Signed-off-by: NOr Gerlitz <ogerlitz@mellanox.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

8fcfb4db

net/mlx4: Set steering mode according to device capabilities · c96d97f4

由 Hadar Hen Zion 提交于 7月 05, 2012

Instead of checking the firmware supported steering mode in various
places in the code, add a dedicated field in the mlx4 device capabilities
structure which is written once during the initialization flow and read
across the code.

This also set the grounds for add new steering modes. Currently two modes
are supported, and are named after the ConnectX HW versions A0 and B0.

A0 steering uses mac_index, vlan_index and priority to steer traffic
into pre-defined range of QPs.

B0 steering uses Ethernet L2 hashing rules and is enabled only
if the firmware supports both unicast and multicast B0 steering,

The current steering modes are relevant for Ethernet traffic only,
such that Infiniband steering remains untouched.
Signed-off-by: NHadar Hen Zion <hadarh@mellanox.co.il>
Signed-off-by: NOr Gerlitz <ogerlitz@mellanox.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

c96d97f4

06 7月, 2012 1 次提交

ipv4: Avoid overhead when no custom FIB rules are installed. · f4530fa5

由 David S. Miller 提交于 7月 05, 2012

If the user hasn't actually installed any custom rules, or fiddled
with the default ones, don't go through the whole FIB rules layer.

It's just pure overhead.

Instead do what we do with CONFIG_IP_MULTIPLE_TABLES disabled, check
the individual tables by hand, one by one.

Also, move fib_num_tclassid_users into the ipv4 network namespace.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

f4530fa5

05 7月, 2012 10 次提交

net-next: Add netif_get_num_default_rss_queues · 16917b87

由 Yuval Mintz 提交于 7月 01, 2012

Most multi-queue networking driver consider the number of online cpus when
configuring RSS queues.
This patch adds a wrapper to the number of cpus, setting an upper limit on the
number of cpus a driver should consider (by default) when allocating resources
for his queues.
Signed-off-by: NYuval Mintz <yuvalmin@broadcom.com>
Signed-off-by: NEilon Greenstein <eilong@broadcom.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

16917b87

D
net: Kill dst->_neighbour, accessors, and final uses. · 36bdbcae
由 David S. Miller 提交于 7月 02, 2012
```
No longer used.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
36bdbcae

ipv6: Store route neighbour in rt6_info struct. · 97cac082

由 David S. Miller 提交于 7月 02, 2012

This makes for a simplified conversion away from dst_get_neighbour*().

All code outside of ipv6 will use neigh lookups via dst_neigh_lookup*().
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

97cac082

D
net: Pass neighbours and dest address into NETEVENT_REDIRECT events. · 1d248b1c
由 David S. Miller 提交于 7月 03, 2012
```
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
1d248b1c

decnet: Use neighbours privately in dn_route struct. · fccd7d5c

由 David S. Miller 提交于 7月 02, 2012

This allows an easy conversion away from dst_get_neighbour*().
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

fccd7d5c

net: Add optional SKB arg to dst_ops->neigh_lookup(). · f894cbf8

由 David S. Miller 提交于 7月 02, 2012

Causes the handler to use the daddr in the ipv4/ipv6 header when
the route gateway is unspecified (local subnet).
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

f894cbf8

net: Do delayed neigh confirmation. · 5110effe

由 David S. Miller 提交于 7月 02, 2012

When a dst_confirm() happens, mark the confirmation as pending in the
dst.  Then on the next packet out, when we have the neigh in-hand, do
the update.

This removes the dependency in dst_confirm() of dst's having an
attached neigh.

While we're here, remove the explicit 'dst' NULL check, all except 2
or 3 call sites ensure it's not NULL.  So just fix those cases up.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

5110effe

ipv4: Make neigh lookups directly in output packet path. · a263b309

由 David S. Miller 提交于 7月 02, 2012

Do not use the dst cached neigh, we'll be getting rid of that.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

a263b309

netfilter: nfnetlink_queue: do not allow to set unsupported flag bits · 46ba5a25

由 Krishna Kumar 提交于 6月 27, 2012

Allow setting of only supported flag bits in queue->flags.
Signed-off-by: NKrishna Kumar <krkumar2@in.ibm.com>
Signed-off-by: NPablo Neira Ayuso <pablo@netfilter.org>

46ba5a25

netfilter: nf_conntrack: generalize nf_ct_l4proto_net · 08911475

由 Pablo Neira Ayuso 提交于 6月 29, 2012

This patch generalizes nf_ct_l4proto_net by splitting it into chunks and
moving the corresponding protocol part to where it really belongs to.

To clarify, note that we follow two different approaches to support per-net
depending if it's built-in or run-time loadable protocol tracker.
Signed-off-by: NPablo Neira Ayuso <pablo@netfilter.org>
Acked-by: NGao feng <gaofeng@cn.fujitsu.com>

08911475

04 7月, 2012 1 次提交

net: em_canid: Ematch rule to match CAN frames according to their identifiers · f057bbb6

由 Rostislav Lisovy 提交于 7月 04, 2012

This ematch makes it possible to classify CAN frames (AF_CAN) according
to their identifiers. This functionality can not be easily achieved with
existing classifiers, such as u32, because CAN identifier is always stored
in native endianness, whereas u32 expects Network byte order.
Signed-off-by: NRostislav Lisovy <lisovy@gmail.com>
Signed-off-by: NOliver Hartkopp <socketcan@hartkopp.net>
Signed-off-by: NMarc Kleine-Budde <mkl@pengutronix.de>

f057bbb6

01 7月, 2012 3 次提交

phy: add the EEE support and the way to access to the MMD registers. · a59a4d19

由 Giuseppe CAVALLARO 提交于 6月 27, 2012

This patch adds the support for the Energy-Efficient Ethernet (EEE)
to the Physical Abstraction Layer.
To support the EEE we have to access to the MMD registers 3.20 and
7.60/61. So two new functions have been added to read/write the MMD
registers (clause 45).

An Ethernet driver (I tested the stmmac) can invoke the phy_init_eee to properly
check if the EEE is supported by the PHYs and it can also set the clock
stop enable bit in the 3.0 register.
The phy_get_eee_err can be used for reporting the number of time where
the PHY failed to complete its normal wake sequence.

In the end, this patch also adds the EEE ethtool support implementing:
 o phy_ethtool_set_eee
 o phy_ethtool_get_eee

v1: initial patch
v2: fixed some errors especially on naming convention
v3: renamed again the mmd read/write functions thank to Ben's feedback
v4: moved file to phy.c and added the ethtool support.
v5: fixed phy_adv_to_eee, phy_eee_to_supported, phy_eee_to_adv return
    values according to ethtool API (thanks to Ben's feedback).
    Renamed some macros to avoid too long names.
v6: fixed kernel-doc comments to be properly parsed.
    Fixed the phy_init_eee function: we need to check which link mode
    was autonegotiated and then the corresponding bits in 7.60 and 7.61
    registers.
v7: reviewed the way to get the negotiated settings.
v8: fixed a problem in the phy_init_eee return value erroneously added
    when included the phy_read_status call.
v9: do not remove the MDIO_AN_EEE_ADV_100TX and MDIO_AN_EEE_ADV_1000T
    and fixed the eee_{cap,lp,adv} declaration as "int" instead of u16.
Signed-off-by: NGiuseppe Cavallaro <peppe.cavallaro@st.com>
Reviewed-by: NBen Hutchings <bhutchings@solarflare.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

a59a4d19

sctp: be more restrictive in transport selection on bundled sacks · 4244854d

由 Neil Horman 提交于 6月 30, 2012

It was noticed recently that when we send data on a transport, its possible that
we might bundle a sack that arrived on a different transport. While this isn't
a major problem, it does go against the SHOULD requirement in section 6.4 of RFC
2960:

An endpoint SHOULD transmit reply chunks (e.g., SACK, HEARTBEAT ACK,
etc.) to the same destination transport address from which it
received the DATA or control chunk to which it is replying. This
rule should also be followed if the endpoint is bundling DATA chunks
together with the reply chunk.

This patch seeks to correct that. It restricts the bundling of sack operations
to only those transports which have moved the ctsn of the association forward
since the last sack. By doing this we guarantee that we only bundle outbound
saks on a transport that has received a chunk since the last sack. This brings
us into stricter compliance with the RFC.

Vlad had initially suggested that we strictly allow only sack bundling on the
transport that last moved the ctsn forward. While this makes sense, I was
concerned that doing so prevented us from bundling in the case where we had
received chunks that moved the ctsn on multiple transports. In those cases, the
RFC allows us to select any of the transports having received chunks to bundle
the sack on. so I've modified the approach to allow for that, by adding a state
variable to each transport that tracks weather it has moved the ctsn since the
last sack. This I think keeps our behavior (and performance), close enough to
our current profile that I think we can do this without a sysctl knob to
enable/disable it.
Signed-off-by: NNeil Horman <nhorman@tuxdriver.com>
CC: Vlad Yaseivch <vyasevich@gmail.com>
CC: David S. Miller <davem@davemloft.net>
CC: linux-sctp@vger.kernel.org
Reported-by: NMichele Baldessari <michele@redhat.com>
Reported-by: Nsorin serban <sserban@redhat.com>
Acked-by: NVlad Yasevich <vyasevich@gmail.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

4244854d

linux/irq.h: fix kernel-doc warning · 87fac288

由 Randy Dunlap 提交于 6月 30, 2012

Fix kernel-doc warning.  This struct member was removed in commit
87568264 ("irq: Remove irq_chip->release()") so remove its
associated kernel-doc entry also.

  Warning(include/linux/irq.h:338): Excess struct/union/enum/typedef member 'release' description in 'irq_chip'
Signed-off-by: NRandy Dunlap <rdunlap@xenotime.net>
Cc: Richard Weinberger <richard@nod.at>
Cc: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

87fac288

30 6月, 2012 2 次提交

net: introduce new priv_flag indicating iface capable of change mac when running · bb35f671

由 Jiri Pirko 提交于 6月 29, 2012

Introduce IFF_LIVE_ADDR_CHANGE priv_flag and use it to disable
netif_running() check in eth_mac_addr()
Signed-off-by: NJiri Pirko <jpirko@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

bb35f671

netlink: add nlk->netlink_bind hook for module auto-loading · 03292745

由 Pablo Neira Ayuso 提交于 6月 29, 2012

This patch adds a hook in the binding path of netlink.

This is used by ctnetlink to allow module autoloading for the case
in which one user executes:

 conntrack -E

So far, this resulted in nfnetlink loaded, but not
nf_conntrack_netlink.

I have received in the past many complains on this behaviour.
Signed-off-by: NPablo Neira Ayuso <pablo@netfilter.org>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

03292745